← Search

Alexis Palmer

12 accepted papers

2025

Boosting the Capabilities of Compact Models in Low-Data Contexts with Large Language Models and Retrieval-Augmented Generation

COLING 2025main

The data and compute requirements of current language modeling technology pose challenges for the processing and analysis of low-resource languages. Declarative linguistic knowledge has the potential to partially bridge this data scarcity gap by providing models with useful inductive bias in the for…

Cited by 3SourcePDFScholar
2025

From Priest to Doctor: Domain Adaptation for Low-Resource Neural Machine Translation

COLING 2025main

Many of the world’s languages have insufficient data to train high-performing general neural machine translation (NMT) models, let alone domain-specific models, and often the only available parallel data are small amounts of religious texts. Hence, domain adaptation (DA) is a crucial issue faced by…

2025

Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentation

EMNLP 2025

Computational morphology has the potential to support language documentation through tasks like morphological segmentation and the generation of Interlinear Glossed Text (IGT). However, our research outputs have seen limited use in real-world language documentation settings. This position paper situ

Cited by 0SourcePDFScholar
2025

Understanding the Gap: an Analysis of Research Collaborations in NLP and Language Documentation

ACL 2025finding

Despite over 20 years of NLP work explicitly intended for application in language documentation (LD), practical use of this work remains vanishingly scarce. This issue has been noted and discussed over the past 10 years, but without the benefit of data to inform the discourse.To address this lack in…

2024

Bootstrapping UMR Annotations for Arapaho from Language Documentation Resources

COLING 2024main

Uniform Meaning Representation (UMR) is a semantic labeling system in the AMR family designed to be uniformly applicable to typologically diverse languages. The UMR labeling system is quite thorough and can be time-consuming to execute, especially if annotators are starting from scratch. In this pap…

Cited by 3SourcePDFScholar
2024

Building a Broad Infrastructure for Uniform Meaning Representations

COLING 2024main

This paper reports the first release of the UMR (Uniform Meaning Representation) data set. UMR is a graph-based meaning representation formalism consisting of a sentence-level graph and a document-level graph. The sentence-level graph represents predicate-argument structures, named entities, word se…

Cited by 9SourcePDFScholar
2024

GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text

EMNLP 2024main

Language documentation projects often involve the creation of annotated text in a format such as interlinear glossed text (IGT), which captures fine-grained morphosyntactic analyses in a morpheme-by-morpheme format. However, there are few existing resources providing large amounts of standardized, e…

Cited by 16SourcePDFScholar
2024

TAMS: Translation-Assisted Morphological Segmentation

ACL 2024long

Canonical morphological segmentation is the process of analyzing words into the standard (aka underlying) forms of their constituent morphemes.This is a core task in endangered language documentation, and NLP systems have the potential to dramatically speed up this process. In typical language docum…

2022

AmericasNLI: Evaluating Zero-shot Natural Language Understanding of Pretrained Multilingual Models in Truly Low-resource Languages

ACL 2022long

Pretrained multilingual models are able to perform cross-lingual transfer in a zero-shot setting, even for languages unseen during pretraining. However, prior work evaluating performance on unseen languages has largely been limited to low-level, syntactic tasks, and it remains unclear if zero-shot l…