← Search

Thomas Friedrichs

3 accepted papers

2022

FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation

EMNLP 2022main

Recent model-based reference-free metrics for open-domain dialogue evaluation exhibit promising correlations with human judgment. However, they either perform turn-level evaluation or look at a single dialogue quality dimension. One would expect a good evaluation metric to assess multiple quality di…

2022

MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue Evaluation

AAAI 2022technical

Chatbots are designed to carry out human-like conversations across different domains, such as general chit-chat, knowledge exchange, and persona-grounded conversations. To measure the quality of such conversational agents, a dialogue evaluator is expected to conduct assessment across domains as well…

2021

DynaEval: Unifying Turn and Dialogue Level Evaluation

ACL 2021long

A dialogue is essentially a multi-turn interaction among interlocutors. Effective evaluation metrics should reflect the dynamics of such interaction. Existing automatic metrics are focused very much on the turn-level quality, while ignoring such dynamics. To this end, we propose DynaEval, a unified…