← Search

Ann Yuan

6 accepted papers

2024

ConstitutionalExperts: Training a Mixture of Principle-based Prompts

ACL 2024short

Large language models (LLMs) are highly capable at a variety of tasks given the right prompt, but writing one is still a difficult and tedious process. In this work, we introduce ConstitutionalExperts, a method for learning a prompt consisting of constitutional principles (i.e. rules), given a train…

Cited by 6SourcePDFScholar
2024

Who's asking? User personas and the mechanics of latent misalignment

NeurIPS 2024spotlight

Studies show that safety-tuned models may nevertheless divulge harmful information. In this work, we show that whether they do so depends significantly on who they are talking to, which we refer to as *user persona*. In fact, we find manipulating user persona to be more effective for eliciting harmf…

Cited by 5SourcePDFScholar
2022

A Recipe for Arbitrary Text Style Transfer with Large Language Models

ACL 2022short

In this paper, we leverage large language models (LLMs) to perform zero-shot text style transfer. We present a prompting method that we call augmented zero-shot learning, which frames style transfer as a sentence rewriting task and requires only a natural language instruction, without model fine-tun…

Cited by 186SourcePDFScholar
2022

The Case for a Single Model that can Both Generate Continuations and Fill-in-the-Blank

NAACL 2022findings

The task of inserting text into a specified position in a passage, known as fill in the blank (FitB), is useful for a variety of applications where writers interact with a natural language generation (NLG) system to craft text. While previous work has tackled this problem with models trained specifi…

Cited by 3SourcePDFScholar
2021

SynthBio: A Case Study in Faster Curation of Text Datasets

NeurIPS 2021poster

NLP researchers need more, higher-quality text datasets. Human-labeled datasets are expensive to collect, while datasets collected via automatic retrieval from the web such as WikiBio [Lebret 2016] are noisy and can include undesired biases. Moreover, data sourced from the web is often included in d…

Cited by 14SourceScholar
2019

Visualizing and Measuring the Geometry of BERT

NeurIPS 2019poster

Transformer architectures show significant promise for natural language processing. Given that a single pretrained model can be fine-tuned to perform well on many different tasks, these networks appear to extract generally useful linguistic features. A natural question is how such networks represent…

Cited by 512SourcePDFScholar