Non-intrusive Speech Quality Assessment for Super-wideband Speech Communication Networks
Gabriel Mittag, Sebastian Möller
Abstract
The quality of speech communication networks has recently improved significantly by extending the available audio bandwidth from narrowband, firstly to wideband, and then to super-wideband. This bandwidth extension marks the end of the typically muffled sound we know from plain old telephone services. Another reason for increased speech quality is the fully digitally packet-based transmission. However, so far, no speech quality prediction model is able to estimate super-wideband quality without a clean reference signal. In this paper, we present a non-intrusive speech quality assessment model NISQA, which - in contrast to current state-of-the-art models - can predict the quality of super-wideband speech transmission. Furthermore, it is able to accurately predict the quality impact of packet loss concealment of modern codecs, such as Opus and EVS. The model uses a novel approach, where a CNN firstly estimates the per-frame quality, and subsequently, an RNN aggregates the per-frame values over time, to estimate the overall speech quality. Averaged over a comprehensive test set, the model achieves an RMSE*3rd of 0.29 with subjective MOS.
BibTeX
@inproceedings{icassp2019_nonintrusivespee,
title = {Non-intrusive Speech Quality Assessment for Super-wideband Speech Communication Networks},
author = {Gabriel Mittag and Sebastian Möller},
booktitle = {ICASSP 2019},
year = {2019}
}