← Search

Arya D. McCarthy

9 accepted papers

2025

LawInstruct: A Resource for Studying Language Model Adaptation to the Legal Domain

NAACL 2025findings

Instruction tuning is an important step in making language models useful for direct user interaction. However, the legal domain is underrepresented in typical instruction datasets (e.g., only 10 out of 1600+ tasks in Super-NaturalInstructions). To study whether instruction tuning on legal datasets i…

2022

A Major Obstacle for NLP Research: Let’s Talk about Time Allocation!

EMNLP 2022main

The field of natural language processing (NLP) has grown over the last few years: conferences have become larger, we have published an incredible amount of papers, and state-of-the-art research has been implemented in a large variety of customer-facing products. However, this paper argues that we ha…

Cited by 2SourcePDFScholar
2022

Deciphering and Characterizing Out-of-Vocabulary Words for Morphologically Rich Languages

COLING 2022main

This paper presents a detailed foundational empirical case study of the nature of out-of-vocabulary words encountered in modern text in a moderate-resource language such as Bulgarian, and a multi-faceted distributional analysis of the underlying word-formation processes that can aid in their composi…

2022

Pre-Trained Multilingual Sequence-to-Sequence Models: A Hope for Low-Resource Language Translation?

ACL 2022findings

What can pre-trained multilingual sequence-to-sequence models like mBART contribute to translating low-resource languages? We conduct a thorough empirical experiment in 10 languages to ascertain this, considering five factors: (1) the amount of fine-tuning data, (2) the noise in the fine-tuning data…

Cited by 33SourcePDFScholar
2021

Jump-Starting Item Parameters for Adaptive Language Tests

EMNLP 2021main

A challenge in designing high-stakes language assessments is calibrating the test item difficulties, either a priori or from limited pilot test data. While prior work has addressed ‘cold start’ estimation of item difficulties without piloting, we devise a multi-task generalized linear model with BER…

Cited by 42SourcePDFScholar
2020

Neural Transduction for Multilingual Lexical Translation

COLING 2020main

We present a method for completing multilingual translation dictionaries. Our probabilistic approach can synthesize new word forms, allowing it to operate in settings where correct translations have not been observed in text (cf. cross-lingual embeddings). In addition, we propose an approximate Maxi…

Cited by 7SourcePDFScholar
2020

SkinAugment: Auto-Encoding Speaker Conversions for Automatic Speech Translation

ICASSP 2020accepted

We propose autoencoding speaker conversion for training data augmentation in automatic speech translation. This technique directly transforms an audio sequence, resulting in audio thesized to resemble another speaker's voice. Our method compares favorably to SpecAugment on English-French and English…

Cited by 0SourceScholar