← Search

George Foster

7 accepted papers

2024

Finding Replicable Human Evaluations via Stable Ranking Probability

NAACL 2024long

Reliable human evaluation is critical to the development of successful natural language generation models, but achieving it is notoriously difficult. Stability is a crucial requirement when ranking systems by quality: consistent ranking of systems across repeated evaluations is not just desirable, b…

2023

Prompting PaLM for Translation: Assessing Strategies and Performance

ACL 2023long

Large language models (LLMs) that have been trained on multilingual but not parallel text exhibit a remarkable ability to translate between languages. We probe this ability in an in-depth study of the pathways language model (PaLM), which has demonstrated the strongest machine translation (MT) perfo…

Cited by 170SourcePDFScholar
2023

Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM’s Translation Capability

ACL 2023long

Large, multilingual language models exhibit surprisingly good zero- or few-shot machine translation capabilities, despite having never seen the intentionally-included translation examples provided to typical neural translation systems. We investigate the role of incidental bilingualism—the unintenti…

Cited by 59SourcePDFScholar
2023

The Unreasonable Effectiveness of Few-shot Learning for Machine Translation

ICML 2023poster

We demonstrate the potential of few-shot translation systems, trained with unpaired language data, for both high and low-resource language pairs. We show that with only 5 examples of high-quality translation data shown at inference, a transformer decoder-only model trained solely with self-supervise…

Cited by 85SourcePDFScholar
2023

Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration

EMNLP 2023long main

Kendall's tau is frequently used to meta-evaluate how well machine translation (MT) evaluation metrics score individual translations. Its focus on pairwise score comparisons is intuitive but raises the question of how ties should be handled, a gray area that has motivated different variants in the…

Cited by 0SourcecodeScholar
2022

A Natural Diet: Towards Improving Naturalness of Machine Translation Output

ACL 2022findings

Machine translation (MT) evaluation often focuses on accuracy and fluency, without paying much attention to translation style. This means that, even when considered accurate and fluent, MT output can still sound less natural than high quality human translations or text originally written in the targ…

Cited by 18SourcePDFScholar
2021

Assessing Reference-Free Peer Evaluation for Machine Translation

NAACL 2021long

Reference-free evaluation has the potential to make machine translation evaluation substantially more scalable, allowing us to pivot easily to new languages or domains. It has been recently shown that the probabilities given by a large, multilingual model can achieve state of the art results when us…

Cited by 22SourcePDFScholar