← Search

Daniel Deutsch

15 accepted papers

2025

Don’t Sweat the Small Stuff: Segment-Level Meta-Evaluation Based on Pairwise Difference Correlation

EMNLP 2025

This paper introduces Pairwise Difference Pearson (PDP), a novel segment-level meta-evaluation metric for Machine Translation (MT) that addresses limitations in previous Pearson’s 𝜌 -based and Kendall’s 𝜏 -based meta-evaluation approaches. PDP is a correlation-based metric that utilizes pairwise dif

Cited by 0SourcePDFScholar
2025

Enhancing Human Evaluation in Machine Translation with Comparative Judgement

ACL 2025long

Human evaluation is crucial for assessing rapidly evolving language models but is influenced by annotator proficiency and task design. This study explores the integration of comparative judgment into human annotation for machine translation (MT) and evaluates three annotation setups—point-wise Multi…

Cited by 0SourcePDFScholar
2025

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set

ICML 2025poster

As LLMs continue to become more powerful and versatile, human evaluation has become intractable at scale and reliance on automatic metrics has become the norm. Recently, it has been shown that LLMs are themselves state-of-the-art evaluators for many tasks. These *Autoraters* are typically designed s…

Cited by 0SourcePDFScholar
2025

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination’s Impact on Machine Translation

ICML 2025poster

Data contamination—the accidental consumption of evaluation examples within the pre-training data—can undermine the validity of evaluation benchmarks. In this paper, we present a rigorous analysis of the effects of contamination on language models at 1B and 8B scales on the machine translation task.…

Cited by 0SourcePDFScholar
2025

SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?

EMNLP 2025

Evaluating machine translation (MT) quality for under-resourced African languages remains a significant challenge, as existing metrics often suffer from limited language coverage and poor performance in low-resource settings. While recent efforts, such as AfriCOMET, have addressed some of the issues

Cited by 0SourcePDFScholar
2025

WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects

ACL 2025finding

As large language models (LLM) become more and more capable in languages other than English, it is important to collect benchmark datasets in order to evaluate their multilingual performance, including on tasks like machine translation (MT). In this work, we extend the WMT24 dataset to cover 55 lang…

Cited by 0SourcePDFScholar
2024

Finding Replicable Human Evaluations via Stable Ranking Probability

NAACL 2024long

Reliable human evaluation is critical to the development of successful natural language generation models, but achieving it is notoriously difficult. Stability is a crucial requirement when ranking systems by quality: consistent ranking of systems across repeated evaluations is not just desirable, b…

2024

LLMRefine: Pinpointing and Refining Large Language Models via Fine-Grained Actionable Feedback

NAACL 2024findings

Recent large language models (LLM) areleveraging human feedback to improve theirgeneration quality. However, human feedbackis costly to obtain, especially during inference.In this work, we propose LLMRefine, aninference time optimization method to refineLLM’s output. The core idea is to usea learned…

Cited by 21SourcePDFScholar
2024

On the Role of Summary Content Units in Text Summarization Evaluation

NAACL 2024short

At the heart of the Pyramid evaluation method for text summarization lie human written summary content units (SCUs). These SCUs areconcise sentences that decompose a summary into small facts. Such SCUs can be used to judge the quality of a candidate summary, possibly partially automated via natural…

2023

A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for Summarization

ACL 2023long

To prevent the costly and inefficient use of resources on low-quality annotations, we want a method for creating a pool of dependable annotators who can effectively complete difficult tasks, such as evaluating automatic summarization. Thus, we investigate the recruitment of high-quality Amazon Mecha…

Cited by 11SourcePDFScholar
2023

Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration

EMNLP 2023long main

Kendall's tau is frequently used to meta-evaluate how well machine translation (MT) evaluation metrics score individual translations. Its focus on pairwise score comparisons is intuitive but raises the question of how ties should be handled, a gray area that has motivated different variants in the…

Cited by 0SourcecodeScholar
2022

Benchmarking Answer Verification Methods for Question Answering-Based Summarization Evaluation Metrics

ACL 2022findings

Question answering-based summarization evaluation metrics must automatically determine whether the QA model’s prediction is correct or not, a task known as answer verification. In this work, we benchmark the lexical answer verification methods which have been used by current QA-based metrics as well…

Cited by 7SourcePDFScholar
2022

Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics

NAACL 2022long

How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which the definition of the system-level correlation is inconsistent with how metrics are used to evaluate systems in practice a…

Cited by 41SourcePDFScholar
2020

Is Killed More Significant than Fled? A Contextual Model for Salient Event Detection

COLING 2020main

Identifying the key events in a document is critical to holistically understanding its important information. Although measuring the salience of events is highly contextual, most previous work has used a limited representation of events that omits essential information. In this work, we propose a hi…

Cited by 12SourcePDFScholar