EMNLP 2023long findings0 citations

FFAEval: Evaluating Dialogue System via Free-For-All Ranking

Zeyao Ma, Zijun Yao, Jing Zhang, Jifan Yu, Xiaohan Zhang, Juanzi Li, Jie Tang

Abstract

Evaluating open-domain dialogue systems is currently an open question. Automatic evaluation metrics have shown poor correlation with human assessment in dialogue generation tasks. Human evaluation, which involves annotators for multi-dimension scoring, is trustworthy but time-consuming. In this work, we propose FFAEval, a reliable and efficient human evaluation framework using Free-For-All ranking approach. By sharing the dialogue history, the framework enables annotators to converse with multiple dialogue systems simultaneously in a single-blind, multi-turn manner. The subsequent free-for-all allows annotators to select the most favourable model in each turn from among all the participating dialogue systems. The final performance of each model is represented by calculating the TrueSkill score derived from the free-for-all competition. Our empirical study on English and Chinese dialogue systems demonstrates that FFAEval achieves a strong correlation with score-based human assessment compared to existing evaluation methods. We further prove the efficiency and stability of our framework in additional experiments. The source code and data are available on Github.

Human EvaluationDialogue System Evaluation
BibTeX
@inproceedings{
ma2023ffaeval,
title={{FFAE}val: Evaluating Dialogue System via Free-For-All Ranking},
author={Zeyao Ma and Zijun Yao and Jing Zhang and Jifan Yu and Xiaohan Zhang and Juanzi Li and Jie Tang},
booktitle={The 2023 Conference on Empirical Methods in Natural Language Processing},
year={2023},
url={https://openreview.net/forum?id=0bderX6zwr}
}
FFAEval: Evaluating Dialogue System via Free-For-All Ranking · EMNLP 2023