← Search

Justin Vasselli

6 accepted papers

2025

Beyond Film Subtitles: Is YouTube the Best Approximation of Spoken Vocabulary?

COLING 2025main

Word frequency is a key variable in psycholinguistics, useful for modeling human familiarity with words even in the era of large language models (LLMs). Frequency in film subtitles has proved to be a particularly good approximation of everyday language exposure. For many languages, however, film sub…

2025

CoAM: Corpus of All-Type Multiword Expressions

ACL 2025long

Multiword expressions (MWEs) refer to idiomatic sequences of multiple words.MWE identification, i.e., detecting MWEs in text, can play a key role in downstream tasks such as machine translation, but existing datasets for the task are inconsistently annotated, limited to a single type of MWE, or limi…

2025

Dictionaries to the Rescue: Cross-Lingual Vocabulary Transfer for Low-Resource Languages Using Bilingual Dictionaries

ACL 2025finding

Cross-lingual vocabulary transfer plays a promising role in adapting pre-trained language models to new languages, including low-resource languages.Existing approaches that utilize monolingual or parallel corpora face challenges when applied to languages with limited resources.In this work, we propo…

2025

How to Make the Most of LLMs’ Grammatical Knowledge for Acceptability Judgments

NAACL 2025long

The grammatical knowledge of language models (LMs) is often measured using a benchmark of linguistic minimal pairs, where LMs are presented with a pair of acceptable and unacceptable sentences and required to judge which is more acceptable. Conventional approaches compare sentence probabilities dire…

2025

Measuring the Robustness of Reference-Free Dialogue Evaluation Systems

COLING 2025main

Advancements in dialogue systems powered by large language models (LLMs) have outpaced the development of reliable evaluation metrics, particularly for diverse and creative responses. We present a benchmark for evaluating the robustness of reference-free dialogue metrics against four categories of a…

2025

Multilingual Dialogue Generation and Localization with Dialogue Act Scripting

EMNLP 2025

Non-English dialogue datasets are scarce, and models are often trained or evaluated on translations of English-language dialogues, an approach which can introduce artifacts that reduce their naturalness and cultural appropriateness. This work proposes Dialogue Act Script (DAS), a structured framewor

Cited by 0SourcePDFScholar