Predicting dialogue success, naturalness, and length with acoustic features
Alexandros Papangelis, Margarita Kotti, Yannis Stylianou
Abstract
Statistical methods for Spoken Dialogue Systems have been shown to reduce the cost of development, while successfully handling a variety of applications. However, such systems are usually trained with simulated users or paid subjects in controlled settings. While this may be sufficient to jump-start learning in the various sub-components, learning is very much dependent on the complete knowledge that we have about the interaction. Relatively few works have focused on this problem, and we here propose to extract low-level audio descriptors and use them as input to various classifiers, namely support vector machines, Gaussian process regressors, and random forests, to predict metrics that are constituents of user satisfaction from acoustic features. While our approach is not directly comparable to the current state of the art, results show that models using the proposed feature set outperform models that use state of the art features extracted from the belief state.
BibTeX
@inproceedings{icassp2017_predictingdialog,
title = {Predicting dialogue success, naturalness, and length with acoustic features},
author = {Alexandros Papangelis and Margarita Kotti and Yannis Stylianou},
booktitle = {ICASSP 2017},
year = {2017}
}