← Search

Pius Von Däniken

5 accepted papers

2024

Favi-Score: A Measure for Favoritism in Automated Preference Ratings for Generative AI Evaluation

ACL 2024long

Generative AI systems have become ubiquitous for all kinds of modalities, which makes the issue of the evaluation of such models more pressing. One popular approach is preference ratings, where the generated outputs of different systems are shown to evaluators who choose their preferences. In recent…

2023

Correction of Errors in Preference Ratings from Automated Metrics for Text Generation

ACL 2023findings

A major challenge in the field of Text Generation is evaluation: Human evaluations are cost-intensive, and automated metrics often display considerable disagreements with human judgments. In this paper, we propose to apply automated metrics for Text Generation in a preference-based evaluation protoc…

Cited by 2SourcePDFScholar
2022

On the Effectiveness of Automated Metrics for Text Generation Systems

EMNLP 2022finding

A major challenge in the field of Text Generation is evaluation, because we lack a sound theory that can be leveraged to extract guidelines for evaluation campaigns. In this work, we propose a first step towards such a theory that incorporates different sources of uncertainty, such as imperfect auto…

2022

Probing the Robustness of Trained Metrics for Conversational Dialogue Systems

ACL 2022short

This paper introduces an adversarial method to stress-test trained metrics for the evaluation of conversational dialogue systems. The method leverages Reinforcement Learning to find response strategies that elicit optimal scores from the trained metrics. We apply our method to test recently proposed…