← Search

Alon Lavie

2 accepted papers

2024

Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs

EMNLP 2024finding

Although human evaluation remains the gold standard for open-domain dialogue evaluation, the growing popularity of automated evaluation using Large Language Models (LLMs) has also extended to dialogue. However, most frameworks leverage benchmarks that assess older chatbots on aspects such as fluency…

2023

The Inside Story: Towards Better Understanding of Machine Translation Neural Evaluation Metrics

ACL 2023short

Neural metrics for machine translation evaluation, such as COMET, exhibit significant improvements in their correlation with human judgments, as compared to traditional metrics based on lexical overlap, such as BLEU. Yet, neural metrics are, to a great extent, “black boxes” returning a single senten…