← Search

Benjamin Hsu

6 accepted papers

2024

M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation

NAACL 2024short

Document translation poses a challenge for Neural Machine Translation (NMT) systems. Most document-level NMT systems rely on meticulously curated sentence-level parallel data, assuming flawless extraction of text from documents along with their precise reading order. These systems also tend to disre…

2023

Pseudo-label Training and Model Inertia in Neural Machine Translation

ICLR 2023poster

Like many other machine learning applications, neural machine translation (NMT) benefits from over-parameterized deep neural models. However, these models have been observed to be brittle: NMT model predictions are sensitive to small input changes and can show significant variation across re-trainin…

Cited by 1SourcePDFScholar
2023

RAMP: Retrieval and Attribute-Marking Enhanced Prompting for Attribute-Controlled Translation

ACL 2023short

Attribute-controlled translation (ACT) is a subtask of machine translation that involves controlling stylistic or linguistic attributes (like formality and gender) of translation outputs. While ACT has garnered attention in recent years due to its usefulness in real-world applications, progress in t…

Cited by 6SourcePDFScholar
2022

CoCoA-MT: A Dataset and Benchmark for Contrastive Controlled MT with Application to Formality

NAACL 2022findings

The machine translation (MT) task is typically formulated as that of returning a single translation for an input segment. However, in many cases, multiple different translations are valid and the appropriate translation may depend on the intended target audience, characteristics of the speaker, or e…

2022

Contrastive Representation Learning for Cross-Document Coreference Resolution of Events and Entities

NAACL 2022long

Identifying related entities and events within and across documents is fundamental to natural language understanding. We present an approach to entity and event coreference resolution utilizing contrastive representation learning. Earlier state-of-the-art methods have formulated this problem as a bi…

2022

MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine Translation

EMNLP 2022main

As generic machine translation (MT) quality has improved, the need for targeted benchmarks that explore fine-grained aspects of quality has increased. In particular, gender accuracy in translation can have implications in terms of output fluency, translation accuracy, and ethics. In this paper, we i…