← Search

Megan Ung

6 accepted papers

2025

Arbiters of Ambivalence: Challenges of using LLMs in No-Consensus tasks

ACL 2025finding

The increasing use of LLMs as substitutes for humans in “aligning” LLMs has raised questions about their ability to replicate human judgments and preferences, especially in ambivalent scenarios where humans disagree. This study examines the biases and limitations of LLMs in three roles: answer gener…

Cited by 0SourcePDFScholar
2025

Improving Model Evaluation using SMART Filtering of Benchmark Datasets

NAACL 2025long

One of the most challenging problems facing NLP today is evaluation. Some of the most pressing issues pertain to benchmark saturation, data contamination, and diversity in the quality of test examples. To address these concerns, we propose Selection Methodology for Accurate, Reduced, and Targeted (S…

Cited by 2SourcePDFScholar
2023

Learning New Skills after Deployment: Improving open-domain internet-driven dialogue with human feedback

ACL 2023long

Frozen models trained to mimic static datasets can never improve their performance. Models that can employ internet-retrieval for up-to-date information and obtain feedback from humans during deployment provide the promise of both adapting to new information, and improving their performance. In this…

Cited by 41SourcePDFScholar
2023

ROBBIE: Robust Bias Evaluation of Large Generative Language Models

EMNLP 2023long main

As generative large language models (LLMs) grow more performant and prevalent, we must develop comprehensive enough tools to measure and improve their fairness. Different prompt-based datasets can be used to measure social bias across multiple text domains and demographic axes, meaning that testing…

Cited by 0SourceScholar
2023

Training Models to Generate, Recognize, and Reframe Unhelpful Thoughts

ACL 2023long

Many cognitive approaches to well-being, such as recognizing and reframing unhelpful thoughts, have received considerable empirical support over the past decades, yet still lack truly widespread adoption in self-help format. A barrier to that adoption is a lack of adequately specific and diverse ded…