← Search

Tobias Domhan

4 accepted papers

2025

Same evaluation, more tokens: On the effect of input length for machine translation evaluation using Large Language Models

EMNLP 2025

Accurately evaluating machine-translated text remains a long-standing challenge, particularly for long documents. Recent work has shown that large language models (LLMs) can serve as reliable and interpretable sentence-level translation evaluators via MQM error span annotations. With modern LLMs sup

2024

A Shocking Amount of the Web is Machine Translated: Insights from Multi-Way Parallelism

ACL 2024findings

We show that content on the web is often translated into many languages, and the low quality of these multi-way translations indicates they were likely created using Machine Translation (MT). Multi-way parallel, machine generated content not only dominates the translations in lower resource language…

2022

The Devil is in the Details: On the Pitfalls of Vocabulary Selection in Neural Machine Translation

NAACL 2022long

Vocabulary selection, or lexical shortlisting, is a well-known technique to improve latency of Neural Machine Translation models by constraining the set of allowed output words during inference. The chosen set is typically determined by separately trained alignment model parameters, independent of t…

2021

Improving the Quality Trade-Off for Neural Machine Translation Multi-Domain Adaptation

EMNLP 2021main

Building neural machine translation systems to perform well on a specific target domain is a well-studied problem. Optimizing system performance for multiple, diverse target domains however remains a challenge. We study this problem in an adaptation setting where the goal is to preserve the existing…

Cited by 8SourcePDFScholar