← Search

Amber Afshan

4 accepted papers

2025

SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models

ICASSP 2025accepted

Speaker Diarization (SD) is a crucial component of modern end-to-end ASR pipelines. Traditional SD systems, which are typically audio-based and operate independently of ASR, often introduce speaker errors, particularly during speaker transitions and overlapping speech. Recently, language models incl…

Cited by 0SourceScholar
2021

Bi-APC: Bidirectional Autoregressive Predictive Coding for Unsupervised Pre-Training and its Application to Children's ASR

ICASSP 2021accepted

We present a bidirectional unsupervised model pre-training (UPT) method and apply it to children’s automatic speech recognition (ASR). An obstacle to improving child ASR is the scarcity of child speech databases. A common approach to alleviate this problem is model pre-training using data from adult…

Cited by 0SourceScholar
2019

Target and Non-target Speaker Discrimination by Humans and Machines

ICASSP 2019accepted

The manner in which acoustic features contribute to perceiving speaker identity remains unclear. In an attempt to better understand speaker perception, we investigated human and machine speaker discrimination with utterances shorter than 2 seconds. Sixty-five listeners performed a same vs. different…

Cited by 0SourceScholar
2016

Better acoustic normalization in subject independent acoustic-to-articulatory inversion: Benefit to recognition

ICASSP 2016accepted

In subject independent acoustic-to-articulatory inversion (SII), the training and test subjects are in general different, whereas subject dependent inversion (SDI) uses the same training and test subjects. Thus, acoustic normalization is used to compensate for the mismatch between the training and t…

Cited by 0SourceScholar