Target and Non-target Speaker Discrimination by Humans and Machines
Soo Jin Park, Amber Afshan, Jody Kreiman, Gary Yeung, Abeer Alwan
Abstract
The manner in which acoustic features contribute to perceiving speaker identity remains unclear. In an attempt to better understand speaker perception, we investigated human and machine speaker discrimination with utterances shorter than 2 seconds. Sixty-five listeners performed a same vs. different task. Machine performance was estimated with i-vector/PLDA-based automatic speaker verification systems, one using mel-frequency cepstral coefficients (MFCCs) and the other using voice quality features (VQual2) inspired by a psychoacoustic model of voice quality. Machine performance was measured in terms of the detection and log-likelihood-ratio cost functions. Humans showed higher confidence for correct target decisions compared to correct non-target decisions, suggesting that they rely on different features and/or decision making strategies when identifying a single speaker compared to when distinguishing between speakers. For non-target trials, responses were highly correlated between humans and the VQual2-based system, especially when speakers were perceptually marked. Fusing human responses with an MFCC-based system improved performance over human-only or MFCC-only results, while fusing with the VQual2-based system did not. The study is a step towards understanding human speaker discrimination strategies and suggests that automatic systems might be able to supplement human decisions especially when speakers are marked.
BibTeX
@inproceedings{icassp2019_targetandnontarg,
title = {Target and Non-target Speaker Discrimination by Humans and Machines},
author = {Soo Jin Park and Amber Afshan and Jody Kreiman and Gary Yeung and Abeer Alwan},
booktitle = {ICASSP 2019},
year = {2019}
}