← Search

Daniel Edward Licht

3 accepted papers

2025

BOUQuET : dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation

EMNLP 2025

BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. This dataset is handcrafted in 8 non-English languages (i.e. Egyptian Arabic and Modern Standard Arabic, French, German, Hindi, Indonesian, Mandarin Chinese, Russian, and Spa

Cited by 0SourcePDFScholar
2023

Multilingual Holistic Bias: Extending Descriptors and Patterns to Unveil Demographic Biases in Languages at Scale

EMNLP 2023long main

We introduce a multilingual extension of the HolisticBias dataset, the largest English template-based taxonomy of textual people references: Multilingual HolisticBias. This extension consists of 20,459 sentences in 50 languages distributed across 13 demographic axes. Source sentences are built from…

Cited by 0SourceScholar
2023

Toxicity in Multilingual Machine Translation at Scale

EMNLP 2023long findings

Machine Translation systems can produce different types of errors, some of which are characterized as critical or catastrophic due to the specific negative impact that they can have on users. In this paper we focus on one type of critical error: added toxicity. We evaluate and analyze added toxicity…

Cited by 0SourceScholar