ICASSP 2025accepted0 citations

A Unified Metric for Simultaneous Evaluation of Error Rate and Annotation Cost

Mark Lindsey, Francis Kubala, Richard M. Stern

Abstract

Pattern classification systems have traditionally been trained using a set of labeled training data and subsequently evaluated using different testing data. The cost of labeling the training data is typically substantial. Online Human-In-The-Loop (HITL) algorithms present an alternate approach that enables useful classification for many real-world applications using much less labeled data. These classifiers begin with a very small amount of training data and iteratively improve their performance by labeling a selected small number of utterances manually. Unfortunately, there is no unified evaluation metric that considers both classifier performance and annotation cost, which makes it difficult to evaluate these algorithms objectively. Furthermore, the lack of such a metric restricts the evaluation of online learning algorithms to prequential evaluation (before the classifier is adapted to the newly-labeled evaluation data), which does not realistically reflect the algorithm’s ability to adapt to the data stream in real time. This paper introduces the Interactive Machine Learning Metric (IMLM), a new unified evaluation metric that makes the combination of performance and annotation cost for binary classification tasks far less arbitrary. This metric is well suited for the evaluation of online HITL algorithms and also allows for fair comparison of different algorithms after adapting to the evaluation data. The value and appropriateness of IMLM is demonstrated by evaluating a series of Online Active Learning algorithms on a Spoken Language Verification task.

BibTeX
@inproceedings{icassp2025_aunifiedmetricfo,
  title = {A Unified Metric for Simultaneous Evaluation of Error Rate and Annotation Cost},
  author = {Mark Lindsey and Francis Kubala and Richard M. Stern},
  booktitle = {ICASSP 2025},
  year = {2025}
}