← Search

Katharina Kann

13 accepted papers

2023

An Investigation of Noise in Morphological Inflection

ACL 2023findings

With a growing focus on morphological inflection systems for languages where high-quality data is scarce, training data noise is a serious but so far largely ignored concern. We aim at closing this gap by investigating the types of noise encountered within a pipeline for truly unsupervised morpholog…

2023

Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers

ACL 2023long

In recent years machine translation has become very successful for high-resource language pairs. This has also sparked new interest in research on the automatic translation of low-resource languages, including Indigenous languages. However, the latter are deeply related to the ethnic and cultural gr…

Cited by 28SourcePDFScholar
2023

Mind the Gap between the Application Track and the Real World

ACL 2023short

Recent advances in NLP have led to a rise in inter-disciplinary and application-oriented research. While this demonstrates the growing real-world impact of the field, research papers frequently feature experiments that do not account for the complexities of realistic data and environments. To explor…

Cited by 2SourcePDFScholar
2022

A Comprehensive Comparison of Neural Networks as Cognitive Models of Inflection

EMNLP 2022main

Neural networks have long been at the center of a debate around the cognitive mechanism by which humans process inflectional morphology. This debate has gravitated into NLP by way of the question: Are neural networks a feasible account for human behavior in morphological inflection?We address that q…

Cited by 4SourcePDFScholar
2022

A Major Obstacle for NLP Research: Let’s Talk about Time Allocation!

EMNLP 2022main

The field of natural language processing (NLP) has grown over the last few years: conferences have become larger, we have published an incredible amount of papers, and state-of-the-art research has been implemented in a large variety of customer-facing products. However, this paper argues that we ha…

Cited by 2SourcePDFScholar
2022

AmericasNLI: Evaluating Zero-shot Natural Language Understanding of Pretrained Multilingual Models in Truly Low-resource Languages

ACL 2022long

Pretrained multilingual models are able to perform cross-lingual transfer in a zero-shot setting, even for languages unseen during pretraining. However, prior work evaluating performance on unseen languages has largely been limited to low-level, syntactic tasks, and it remains unclear if zero-shot l…

2022

BPE vs. Morphological Segmentation: A Case Study on Machine Translation of Four Polysynthetic Languages

ACL 2022findings

Morphologically-rich polysynthetic languages present a challenge for NLP systems due to data sparsity, and a common strategy to handle this issue is to apply subword segmentation. We investigate a wide variety of supervised and unsupervised morphological segmentation methods for four polysynthetic l…

Cited by 25SourcePDFScholar
2022

Match the Script, Adapt if Multilingual: Analyzing the Effect of Multilingual Pretraining on Cross-lingual Transferability

ACL 2022long

Pretrained multilingual models enable zero-shot learning even for unseen languages, and that performance can be further improved via adaptation prior to finetuning. However, it is unclear how the number of pretraining languages influences a model’s zero-shot learning for languages unseen during pret…

2022

Morphological Processing of Low-Resource Languages: Where We Are and What’s Next

ACL 2022findings

Automatic morphological processing can aid downstream natural language processing applications, especially for low-resource languages, and assist language documentation efforts for endangered languages. Having long been multilingual, the field of computational morphology is increasingly moving towar…

2021

Don’t Rule Out Monolingual Speakers: A Method For Crowdsourcing Machine Translation Data

ACL 2021short

High-performing machine translation (MT) systems can help overcome language barriers while making it possible for everyone to communicate and use language technologies in the language of their choice. However, such systems require large amounts of parallel sentences for training, and translators can…

Cited by 2SourcePDFScholar
2021

The World of an Octopus: How Reporting Bias Influences a Language Model’s Perception of Color

EMNLP 2021main

Recent work has raised concerns about the inherent limitations of text-only pretraining. In this paper, we first demonstrate that reporting bias, the tendency of people to not state the obvious, is one of the causes of this limitation, and then investigate to what extent multimodal training can miti…