← Search

Dallas Card

11 accepted papers

2025

Coordinating Chaos: A Structured Review of Linguistic Coordination Methodologies

ACL 2025long

Linguistic coordination—a phenomenon where conversation partners end up having similar patterns of language use—has been established across a variety of contexts and for multiple linguistic features. However, the study of language coordination has been accompanied by a diverse and inconsistently app…

Cited by 0SourcePDFScholar
2025

Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus

ACL 2025long

Podcasts provide highly diverse content to a massive listener base through a unique on-demand modality. However, limited data has prevented large-scale computational analysis of the podcast ecosystem. To fill this gap, we introduce a massive dataset of over 1.1M podcast transcripts that is largely c…

2024

You don’t need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments

NAACL 2024long

The versatility of Large Language Models (LLMs) on natural language understanding tasks has made them popular for research in social sciences. To properly understand the properties and innate personas of LLMs, researchers have performed studies that involve using prompts in the form of questions tha…

2023

When it Rains, it Pours: Modeling Media Storms and the News Ecosystem

EMNLP 2023long findings

Most events in the world receive at most brief coverage by the news media. Occasionally, however, an event will trigger a media storm, with voluminous and widespread coverage lasting for weeks instead of days. In this work, we develop and apply a pairwise article similarity model, allowing us to ide…

Cited by 0SourcecodeScholar
2022

Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words

ACL 2022short

Cosine similarity of contextual embeddings is used in many NLP tasks (e.g., QA, IR, MT) and metrics (e.g., BERTScore). Here, we uncover systematic ways in which word similarities estimated by cosine over BERT embeddings are understated and trace this effect to training data frequency. We find that r…

2022

Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection

EMNLP 2022main

Language models increasingly rely on massive web crawls for diverse text data. However, these sources are rife with undesirable content. As such, resources like Wikipedia, books, and news often serve as anchors for automatically selecting web text most suitable for language modeling, a process typic…

Cited by 31SourcePDFScholar
2021

Causal Effects of Linguistic Properties

NAACL 2021long

We consider the problem of using observational data to estimate the causal effects of linguistic properties. For example, does writing a complaint politely lead to a faster response time? How much will a positive product review increase sales? This paper addresses two technical challenges related to…

2021

Expected Validation Performance and Estimation of a Random Variable’s Maximum

EMNLP 2021finding

Research in NLP is often supported by experimental results, and improved reporting of such results can lead to better understanding and more reproducible science. In this paper we analyze three statistical estimators for expected validation performance, a tool used for reporting performance (e.g., a…

Cited by 6SourcePDFScholar