← Search

Zoey Liu

9 accepted papers

2025

What data should I include in my POS tagging training set?

EMNLP 2025

Building an NLP training set for understudied languages, including Indigenous and endangered languages, often faces challenges due to varying degrees of resource limitations in the speaker communities. What are some reasonable approaches for training set construction in these cases? We address this

2024

Are modern neural ASR architectures robust for polysynthetic languages?

EMNLP 2024finding

Automatic speech recognition (ASR) technology is frequently proposed as a means of preservation and documentation of endangered languages, with promising results thus far. Among the endangered languages spoken today, a significant number exhibit complex morphology. The models employed in contemporar…

2024

Enough Is Enough! a Case Study on the Effect of Data Size for Evaluation Using Universal Dependencies

COLING 2024main

When creating a new dataset for evaluation, one of the first considerations is the size of the dataset. If our evaluation data is too small, we risk making unsupported claims based on the results on such data. If, on the other hand, the data is too large, we waste valuable annotation time and costs…

Cited by 0SourcePDFScholar
2024

How Important is a Language Model for Low-resource ASR?

ACL 2024findings

N-gram language models (LMs) are the innovation that first made large-vocabulary continuous automatic speech recognition (ASR) viable. With neural end-to-end ASR architectures, however, LMs have become an afterthought. While the effect on accuracy may be negligible for English and Mandarin, jettison…

Cited by 1SourcePDFScholar
2024

The Effect of Data Partitioning Strategy on Model Generalizability: A Case Study of Morphological Segmentation

NAACL 2024long

Recent work to enhance data partitioning strategies for more realistic model evaluation face challenges in providing a clear optimal choice. This study addresses these challenges, focusing on morphological segmentation and synthesizing limitations related to language diversity, adoption of multiple…

Cited by 0SourcePDFScholar
2023

An (unhelpful) guide to selecting the best ASR architecture for your under-resourced language

ACL 2023short

Advances in deep neural models for automatic speech recognition (ASR) have yielded dramatic improvements in ASR quality for resource-rich languages, with English ASR now achieving word error rates comparable to that of human transcribers. The vast majority of the world’s languages, however, lack the…

Cited by 14SourcePDFScholar
2022

Evaluating the Performance of Transformer-based Language Models for Neuroatypical Language

COLING 2022main

Difficulties with social aspects of language are among the hallmarks of autism spectrum disorder (ASD). These communication differences are thought to contribute to the challenges that adults with ASD experience when seeking employment, underscoring the need for interventions that focus on improving…

Cited by 5SourcePDFScholar
2022

Not always about you: Prioritizing community needs when developing endangered language technology

ACL 2022long

Languages are classified as low-resource when they lack the quantity of data necessary for training statistical and machine learning tools and models. Causes of resource scarcity vary but can include poor access to technology for developing these resources, a relatively small population of speakers,…

Cited by 33SourcePDFScholar