← Search

David R. Mortensen

20 accepted papers

2025

DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models

ACL 2025long

Most of the world’s languages and dialects are low-resource, and lack support in mainstream machine translation (MT) models. However, many of them have a closely-related high-resource language (HRL) neighbor, and differ in linguistically regular ways from it. This underscores the importance of model…

2025

Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment

NAACL 2025long

Allophony refers to the variation in the phonetic realization of a phoneme based on its phonetic environment. Modeling allophones is crucial for atypical pronunciation assessment, which involves distinguishing atypical from typical pronunciations. However, recent phoneme classifier-based approaches…

2025

Programming by Example meets Historical Linguistics: A Large Language Model Based Approach to Sound Law Induction

ACL 2025long

Historical linguists have long written “programs” that convert reconstructed words in an ancestor language into their attested descendants via ordered string rewrite functions (called sound laws) However, writing these programs is time-consuming, motivating the development of automated Sound Law Ind…

2025

ZIPA: A family of efficient models for multilingual phone recognition

ACL 2025long

We present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition. We first curated IPA PACK++, a large-scale multilingual speech corpus with 17,000+ hours of normalized phone transcriptions and a novel evaluation set capturing u…

2024

Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons

COLING 2024main

In this paper, we make a contribution that can be understood from two perspectives: from an NLP perspective, we introduce a small challenge dataset for NLI with large lexical overlap, which minimises the possibility of models discerning entailment solely based on token distinctions, and show that GP…

2024

PWESuite: Phonetic Word Embeddings and Tasks They Facilitate

COLING 2024main

Mapping words into a fixed-dimensional vector space is the backbone of modern NLP. While most word embedding methods successfully encode semantic information, they overlook phonetic information that is crucial for many tasks. We develop three methods that use articulatory features to build phonetica…

2024

Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs

COLING 2024main

Lexical-syntactic flexibility, in the form of conversion (or zero-derivation) is a hallmark of English morphology. In conversion, a word with one part of speech is placed in a non-prototypical context, where it is coerced to behave as if it had a different part of speech. However, while this process…

Cited by 5SourcePDFScholar
2024

Zero-Shot Cross-Lingual NER Using Phonemic Representations for Low-Resource Languages

EMNLP 2024main

Existing zero-shot cross-lingual NER approaches require substantial prior knowledge of the target language, which is impractical for low-resource languages.In this paper, we propose a novel approach to NER using phonemic representation based on the International Phonetic Alphabet (IPA) to bridge the…

2023

Calibrated Seq2seq Models for Efficient and Generalizable Ultra-fine Entity Typing

EMNLP 2023long findings

Ultra-fine entity typing plays a crucial role in information extraction by predicting fine-grained semantic types for entity mentions in text. However, this task poses significant challenges due to the massive number of entity types in the output space. The current state-of-the-art approaches, based…

Cited by 0SourcecodeScholar
2023

Counting the Bugs in ChatGPT's Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language Model

EMNLP 2023long main

Large language models (LLMs) have recently reached an impressive level of linguistic capability, prompting comparisons with human language skills. However, there have been relatively few systematic inquiries into the linguistic capabilities of the latest generation of LLMs, and those studies that do…

Cited by 0SourceScholar
2023

Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models

EMNLP 2023long main

Language models have graduated from being research prototypes to commercialized products offered as web APIs, and recent works have highlighted the multilingual capabilities of these products. The API vendors charge their users based on usage, more specifically on the number of ``tokens'' processed…

Cited by 0SourceScholar
2022

WikiHan: A New Comparative Dataset for Chinese Languages

COLING 2022main

Most comparative datasets of Chinese varieties are not digital; however, Wiktionary includes a wealth of transcriptions of words from these varieties. The usefulness of these data is limited by the fact that they use a wide range of variety-specific romanizations, making data difficult to compare. T…

2021

Evaluating the Morphosyntactic Well-formedness of Generated Texts

EMNLP 2021main

Text generation systems are ubiquitous in natural language processing applications. However, evaluation of these systems remains a challenge, especially in multilingual settings. In this paper, we propose L’AMBRE – a metric to evaluate the morphosyntactic well-formedness of text using its dependency…

2021

Multilingual Phonetic Dataset for Low Resource Speech Recognition

ICASSP 2021accepted

Phone Recognition is one of the most important tasks in the field of multilingual speech recognition, especially for low-resource languages whose orthographies are not available. However, most speech recognition datasets so far only focus on high-resource languages, there are very few datasets avail…

Cited by 0SourceScholar
2020

Universal Phone Recognition with a Multilingual Allophone System

ICASSP 2020accepted

Multilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonemes (sounds that can support lexical contrasts in a particular language) and thei…

Cited by 0SourceScholar