← Search

David Bamman

12 accepted papers

2025

Tell, Don’t Show: Leveraging Language Models’ Abstractive Retellings to Model Literary Themes

ACL 2025finding

Conventional bag-of-words approaches for topic modeling, like latent Dirichlet allocation (LDA), struggle with literary text. Literature challenges lexical methods because narrative language focuses on immersive sensory details instead of abstractive description or exposition: writers are advised to…

2024

AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters

ACL 2024long

Large language models’ (LLMs) abilities are drawn from their pretraining data, and model development begins with data curation. However, decisions around what data is retained or removed during this initial stage are under-scrutinized. In our work, we ground web text, which is a popular pretraining…

2023

Grounding Characters and Places in Narrative Text

ACL 2023long

Tracking characters and locations throughout a story can help improve the understanding of its plot structure. Prior research has analyzed characters and locations from text independently without grounding characters to their locations in narrative time. Here, we address this gap by proposing a new…

2023

Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4

EMNLP 2023long main

In this work, we carry out a data archaeology to infer books that are known to ChatGPT and GPT-4 using a name cloze membership inference query. We find that OpenAI models have memorized a wide collection of copyrighted materials, and that the degree of memorization is tied to the frequency with whic…

Cited by 0SourcecodeScholar
2023

Words as Gatekeepers: Measuring Discipline-specific Terms and Meanings in Scholarly Publications

ACL 2023findings

Scholarly text is often laden with jargon, or specialized language that can facilitate efficient in-group communication within fields but hinder understanding for out-groups. In this work, we develop and validate an interpretable approach for measuring scholarly jargon from text. Expanding the scope…

2022

Discovering Differences in the Representation of People using Contextualized Semantic Axes

EMNLP 2022main

A common paradigm for identifying semantic differences across social and temporal contexts is the use of static word embeddings and their distances. In particular, past work has compared embeddings against “semantic axes” that represent two opposing concepts. We extend this paradigm to BERT embeddin…

2022

Predicting Long-Term Citations from Short-Term Linguistic Influence

EMNLP 2022finding

A standard measure of the influence of a research paper is the number of times it is cited. However, papers may be cited for many reasons, and citation count is not informative about the extent to which a paper affected the content of subsequent publications. We therefore propose a novel method to q…

2019

Learning to Groove with Inverse Sequence Transformations

ICML 2019oral

We explore models for translating abstract musical ideas (scores, rhythms) into expressive performances using seq2seq and recurrent variational information bottleneck (VIB) models. Though seq2seq models usually require painstakingly aligned corpora, we show that it is possible to adapt an approach f…

Cited by 136SourcePDFScholar