← Search

Kathleen McKeown

56 accepted papers

2026

Estimating Tail Risks in Language Model Output Distributions

ICML 2026spotlight

Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these models is increasingly high-stakes. Fortunately, advances in alignment have significantly reduced the likelihood of harmful model outputs. However, when models are qu…

Cited by 0SourceScholar
2026

LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News

ICML 2026poster

Large Language Models (LLMs) with agentic web search capabilities show strong potential for tasks requiring real-time information access and complex fact retrieval, yet evaluating such systems remains challenging. We introduce LiveNewsBench, a rigorous and regularly updated benchmark designed to ass…

Cited by 0SourceScholar
2026

PoSh: Using Scene Graphs to Guide LLMs-as-a-Judge for Detailed Image Descriptions

ICLR 2026poster

While vision-language models (VLMs) have advanced into detailed image description, evaluation remains a challenge. Standard metrics (e.g. CIDEr, SPICE) were designed for short texts and tuned to recognize errors that are now uncommon, such as object misidentification. In contrast, long texts require…

Cited by 0SourcecodeScholar
2025

A General Framework for Inference-time Scaling and Steering of Diffusion Models

ICML 2025poster

Diffusion models have demonstrated remarkable performance in generative modeling, but generating samples with specific desiderata remains challenging. Existing solutions --- such as fine-tuning, best-of-n sampling, and gradient-based guidance --- are expensive, inefficient, or limited in applicabil…

2025

Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions

EMNLP 2025

We introduce and study artificial impressions–patterns in LLMs’ internal representations of prompts that resemble human impressions and stereotypes based on language. We fit linear probes on generated prompts to predict impressions according to the two-dimensional Stereotype Content Model (SCM). Usi

2025

Data Caricatures: On the Representation of African American Language in Pretraining Corpora

ACL 2025long

With a combination of quantitative experiments, human judgments, and qualitative analyses, we evaluate the quantity and quality of African American Language (AAL) representation in 12 predominantly English, open-source pretraining corpora. We specifically focus on the sources, variation, and natural…

Cited by 0SourcePDFScholar
2025

Enhancing Multimodal Affective Analysis with Learned Live Comment Features

AAAI 2025technical

Live comments, also known as Danmaku, are user-generated messages that are synchronized with video content. These comments overlay directly onto streaming videos, capturing viewer emotions and reactions in real-time. While prior work has leveraged live comments in affective analysis, its use has bee…

2025

Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment

EMNLP 2025

Large Language Models (LLMs) are typically trained to reflect a relatively uniform set of values, which limits their applicability to tasks that require understanding of nuanced human perspectives. Recent research has underscored the importance of enabling LLMs to support steerable pluralism — the c

Cited by 0SourcePDFScholar
2025

Guiding LLM Decision-Making with Fairness Reward Models

NeurIPS 2025poster

Large language models are increasingly used to support high-stakes decisions, potentially influencing who is granted bail or receives a loan. Naive chain-of-thought sampling can improve average decision accuracy, but has also been shown to amplify unfair bias. To address this challenge and enable th…

Cited by 0SourcecodeScholar
2025

Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding

EMNLP 2025

Determining faithfulness of a claim to a source document is an important problem across many domains. This task is generally treated as a binary judgment of whether the claim is supported or unsupported in relation to the source. In many cases, though, whether a claim is supported can be ambiguous.

2025

Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution

COLING 2025main

Recent state-of-the-art authorship attribution methods learn authorship representations of text in a latent, uninterpretable space, which hinders their usability in real-world applications. We propose a novel approach for interpreting learned embeddings by identifying representative points in the la…

Cited by 0SourcePDFScholar
2025

Layered Insights: Generalizable Analysis of Human Authorial Style by Leveraging All Transformer Layers

EMNLP 2025

We propose a new approach for the authorship attribution task that leverages the various linguistic representations learned at different layers of pre-trained transformer-based models. We evaluate our approach on two popular authorship attribution models and three evaluation datasets, in in-domain a

Cited by 0SourcePDFScholar
2025

ManiTweet: A New Benchmark for Identifying Manipulation of News on Social Media

COLING 2025main

Considerable advancements have been made to tackle the misrepresentation of information derived from reference articles in the domains of fact-checking and faithful summarization. However, an unaddressed aspect remains - the identification of social media posts that manipulate information within ass…

2025

ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges

EMNLP 2025

Recent progress in large language models (LLMs) has enabled substantial advances in solving mathematical problems. However, existing benchmarks often fail to reflect real-world complexity, which demand open-ended, interdisciplinary reasoning and integration of computational tools. To address this ga

2025

See It from My Perspective: How Language Affects Cultural Bias in Image Understanding

ICLR 2025poster

Vision-language models (VLMs) can respond to queries about images in many languages. However, beyond language, culture affects how we see things. For example, individuals from Western cultures focus more on the central figure in an image while individuals from East Asian cultures attend more to sce…

Cited by 0SourcePDFScholar
2025

StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples

NAACL 2025long

Style representations aim to embed texts with similar writing styles closely and texts with different styles far apart, regardless of content. However, the contrastive triplets often used for training these representations may vary in both style and content, leading to potential content leakage in t…

Cited by 2SourcePDFScholar
2025

Summarization of Opinionated Political Documents with Varied Perspectives

COLING 2025main

Global partisan hostility and polarization has increased, and this polarization is heightened around presidential elections. Models capable of generating accurate summaries of diverse perspectives can help reduce such polarization by exposing users to alternative perspectives. In this work, we intro…

2025

The Law of Knowledge Overshadowing: Towards Understanding, Predicting and Preventing LLM Hallucination

ACL 2025finding

Hallucination is a persistent challenge in large language models (LLMs), where even with rigorous quality control, models often generate distorted facts. This paradox, in which error generation continues despite high-quality training data, calls for a deeper understanding of the underlying LLM mecha…

Cited by 0SourcePDFScholar
2024

Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

ICML 2024spotlight

Large language models (LLMs) are trained to imitate humans to explain human decisions. However, do LLMs explain themselves? Can they help humans build mental models of how LLMs process different inputs? To answer these questions, we propose to evaluate $\textbf{counterfactual simulatability}$ of nat…

Cited by 57SourcePDFScholar
2024

Fair Abstractive Summarization of Diverse Perspectives

NAACL 2024long

People from different social and demographic groups express diverse perspectives and conflicting opinions on a broad set of topics such as product reviews, healthcare, law, and politics. A fair summary should provide a comprehensive coverage of diverse perspectives without underrepresenting certain…

2024

Getting Serious about Humor: Crafting Humor Datasets with Unfunny Large Language Models

ACL 2024short

Humor is a fundamental facet of human cognition and interaction. Yet, despite recent advances in natural language processing, humor detection remains a challenging task that is complicated by the scarcity of datasets that pair humorous texts with similar non-humorous counterparts. We investigate whe…

2024

MASIVE: Open-Ended Affective State Identification in English and Spanish

EMNLP 2024main

In the field of emotion analysis, much NLP research focuses on identifying a limited number of discrete emotion categories, often applied across languages. These basic sets, however, are rarely designed with textual data in mind, and culture, language, and dialect can influence how particular emotio…

2024

ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer

AAAI 2024technical

Textual style transfer is the task of transforming stylistic properties of text while preserving meaning. Target "styles" can be defined in numerous ways, ranging from single attributes (e.g. formality) to authorship (e.g. Shakespeare). Previous unsupervised style-transfer approaches generally rely…

2024

Parallel Structures in Pre-training Data Yield In-Context Learning

ACL 2024long

Pre-trained language models (LMs) are capable of in-context learning (ICL): they can adapt to a task with only a few examples given in the prompt without any parameter update. However, it is unclear where this capability comes from as there is a stark distribution shift between pre-training text and…

Cited by 13SourcePDFScholar
2024

STORYSUMM: Evaluating Faithfulness in Story Summarization

EMNLP 2024main

Human evaluation has been the gold standard for checking faithfulness in abstractive summarization. However, with a challenging source domain like narrative, multiple annotators can agree a summary is faithful, while missing details that are obvious errors only once pointed out. We therefore introdu…

2024

Social Orientation: A New Feature for Dialogue Analysis

COLING 2024main

There are many settings where it is useful to predict and explain the success or failure of a dialogue. Circumplex theory from psychology models the social orientations (e.g., Warm-Agreeable, Arrogant-Calculating) of conversation participants and can be used to predict and explain the outcome of soc…

Cited by 3SourcePDFScholar
2024

TinyStyler: Efficient Few-Shot Text Style Transfer with Authorship Embeddings

EMNLP 2024finding

The goal of text style transfer is to transform the style of texts while preserving their original meaning, often with only a few examples of the target style. Existing style transfer methods generally rely on the few-shot capabilities of large language models or on complex controllable text generat…

2024

TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization

NAACL 2024long

Single document news summarization has seen substantial progress on faithfulness in recent years, driven by research on the evaluation of factual consistency, or hallucinations. We ask whether these advances carry over to other text summarization domains. We propose a new evaluation benchmark on top…

2023

Check-COVID: Fact-Checking COVID-19 News Claims with Scientific Evidence

ACL 2023findings

We present a new fact-checking benchmark, Check-COVID, that requires systems to verify claims about COVID-19 from news using evidence from scientific articles. This approach to fact-checking is particularly challenging as it requires checking internet text written in everyday language against eviden…

2023

Evaluation of African American Language Bias in Natural Language Generation

EMNLP 2023long main

While biases disadvantaging African American Language (AAL) have been uncovered in models for tasks such as speech recognition and toxicity detection, there has been little investigation of these biases for language generation models like ChatGPT. We evaluate how well LLMs understand AAL in comparis…

Cited by 0SourceScholar
2023

Faking Fake News for Real Fake News Detection: Propaganda-Loaded Training Data Generation

ACL 2023long

Despite recent advances in detecting fake news generated by neural models, their results are not readily applicable to effective detection of human-written disinformation. What limits the successful transfer between them is the sizable gap between machine-generated fake news and human-authored ones,…

2023

Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source Learning

ACL 2023long

In this paper, we present a novel approach for data-to-text generation that addresses the limitations of current methods that primarily focus on specific types of structured data. Our proposed method aims to improve performance in multi-task training, zero-shot and few-shot scenarios by providing a…

Cited by 1SourcePDFScholar
2023

Generating EDU Extracts for Plan-Guided Summary Re-Ranking

ACL 2023long

Two-step approaches, in which summary candidates are generated-then-reranked to return a single summary, can improve ROUGE scores over the standard single-step approach. Yet, standard decoding methods (i.e., beam search, nucleus sampling, and diverse beam search) produce candidates with redundant, a…

2023

Improving Long Dialogue Summarization with Semantic Graph Representation

ACL 2023findings

Although Large Language Models (LLMs) are successful in abstractive summarization of short dialogues, summarization of long dialogues remains challenging. To address this challenge, we propose a novel algorithm that processes complete dialogues comprising thousands of tokens into topic-segment-level…

Cited by 11SourcePDFScholar
2023

Learning Interpretable Style Embeddings via Prompting LLMs

EMNLP 2023long findings

Style representation learning builds content-independent representations of author style in text. To date, no large dataset of texts with stylometric annotations on a wide range of style dimensions has been compiled, perhaps because the linguistic expertise to perform such annotation would be prohib…

Cited by 0SourceScholar
2022

Constrained Regeneration for Cross-Lingual Query-Focused Extractive Summarization

COLING 2022main

Query-focused summaries of foreign-language, retrieved documents can help a user understand whether a document is actually relevant to the query term. A standard approach to this problem is to first translate the source documents and then perform extractive summarization to find relevant snippets. H…

Cited by 3SourcePDFScholar
2022

Faithful or Extractive? On Mitigating the Faithfulness-Abstractiveness Trade-off in Abstractive Summarization

ACL 2022long

Despite recent progress in abstractive summarization, systems still suffer from faithfulness errors. While prior work has proposed models that improve faithfulness, it is unclear whether the improvement comes from an increased level of extractiveness of the model outputs as one naive way to improve…

2022

Learning to Revise References for Faithful Summarization

EMNLP 2022finding

In real-world scenarios with naturally occurring datasets, reference summaries are noisy and may contain information that cannot be inferred from the source text. On large news corpora, removing low quality samples has been shown to reduce model hallucinations. Yet, for smaller, and/or noisier corpo…

2022

Mitigating Covertly Unsafe Text within Natural Language Systems

EMNLP 2022finding

An increasingly prevalent problem for intelligent technologies is text safety, as uncontrolled systems may generate recommendations to their users that lead to injury or life-threatening consequences. However, the degree of explicitness of a generated statement that can cause physical harm varies. I…

Cited by 7SourcePDFScholar
2022

Read Top News First: A Document Reordering Approach for Multi-Document News Summarization

ACL 2022findings

A common method for extractive multi-document news summarization is to re-formulate it as a single-document summarization problem by concatenating all documents as a single meta-document. However, this method neglects the relative importance of documents. We propose a simple approach to reorder the…

2022

SafeText: A Benchmark for Exploring Physical Safety in Language Models

EMNLP 2022main

Understanding what constitutes safe text is an important issue in natural language processing and can often prevent the deployment of models deemed harmful and unsafe. One such type of safety that has been scarcely studied is commonsense physical safety, i.e. text that is not explicitly violent and…

2022

Seeded Hierarchical Clustering for Expert-Crafted Taxonomies

EMNLP 2022finding

Practitioners from many disciplines (e.g., political science) use expert-crafted taxonomies to make sense of large, unlabeled corpora. In this work, we study Seeded Hierarchical Clustering (SHC): the task of automatically fitting unlabeled data to such taxonomies using a small set of labeled example…

Cited by 1SourcePDFScholar
2022

Using Structured Content Plans for Fine-grained Syntactic Control in Pretrained Language Model Generation

COLING 2022main

Large pretrained language models offer powerful generation capabilities, but cannot be reliably controlled at a sub-sentential level. We propose to make such fine-grained control possible in pretrained LMs by generating text directly from a semantic representation, Abstract Meaning Representation (A…

2022

What Do Users Care About? Detecting Actionable Insights from User Feedback

NAACL 2022industry

Users often leave feedback on a myriad of aspects of a product which, if leveraged successfully, can help yield useful insights that can lead to further improvements down the line. Detecting actionable insights can be challenging owing to large amounts of data as well as the absence of labels in rea…

Cited by 3SourcePDFScholar
2021

Adversarial Learning for Zero-Shot Stance Detection on Social Media

NAACL 2021long

Stance detection on social media can help to identify and understand slanted news or commentary in everyday life. In this work, we propose a new model for zero-shot stance detection on Twitter that uses adversarial learning to generalize across topics. Our model achieves state-of-the-art performance…

2021

Cross-language Sentence Selection via Data Augmentation and Rationale Training

ACL 2021long

This paper proposes an approach to cross-language sentence selection in a low-resource setting. It uses data augmentation and negative sampling techniques on noisy parallel sentence data to directly learn a cross-lingual embedding-based query relevance model. Results show that this approach performs…

Cited by 11SourcePDFScholar
2021

Emotion-Infused Models for Explainable Psychological Stress Detection

NAACL 2021long

The problem of detecting psychological stress in online posts, and more broadly, of detecting people in distress or in need of help, is a sensitive application for which the ability to interpret models is vital. Here, we present work exploring the use of a semantically related task, emotion detectio…

2021

Improving Factual Consistency of Abstractive Summarization via Question Answering

ACL 2021long

A commonly observed problem with the state-of-the art abstractive summarization models is that the generated summaries can be factually inconsistent with the input documents. The fact that automatic summarization may produce plausible-sounding yet inaccurate summaries is a major concern that limits…

2021

InfoSurgeon: Cross-Media Fine-grained Information Consistency Checking for Fake News Detection

ACL 2021long

To defend against machine-generated fake news, an effective mechanism is urgently needed. We contribute a novel benchmark for fake news detection at the knowledge element level, as well as a solution for this task which incorporates cross-media consistency checking to detect the fine-grained knowled…

2021

Supporting Clustering with Contrastive Learning

NAACL 2021long

Unsupervised clustering aims at discovering the semantic categories of data according to some distance measured in the representation space. However, different categories often overlap with each other in the representation space at the beginning of the learning process, which poses a significant cha…

2021

Timeline Summarization based on Event Graph Compression via Time-Aware Optimal Transport

EMNLP 2021main

Timeline Summarization identifies major events from a news collection and describes them following temporal order, with key dates tagged. Previous methods generally generate summaries separately for each date after they determine the key dates of events. These methods overlook the events’ intra-stru…

2020

Detecting Urgency Status of Crisis Tweets: A Transfer Learning Approach for Low Resource Languages

COLING 2020main

We release an urgency dataset that consists of English tweets relating to natural crises, along with annotations of their corresponding urgency status. Additionally, we release evaluation datasets for two low-resource languages, i.e. Sinhala and Odia, and demonstrate an effective zero-shot transfer…

2020

Event-Guided Denoising for Multilingual Relation Learning

COLING 2020main

General purpose relation extraction has recently seen considerable gains in part due to a massively data-intensive distant supervision technique from Soares et al. (2019) that produces state-of-the-art results across many benchmarks. In this work, we present a methodology for collecting high quality…