← Search

Carolina Zheng

2 accepted papers

2024

Hypothesis Testing the Circuit Hypothesis in LLMs

NeurIPS 2024poster

Large language models (LLMs) demonstrate surprising capabilities, but we do not understand how they are implemented. One hypothesis suggests that these capabilities are primarily executed by small subnetworks within the LLM, known as circuits. But how can we evaluate this hypothesis? In this paper,…

2023

An Invariant Learning Characterization of Controlled Text Generation

ACL 2023long

Controlled generation refers to the problem of creating text that contains stylistic or semantic attributes of interest. Many approaches reduce this problem to training a predictor of the desired attribute. For example, researchers hoping to deploy a large language model to produce non-toxic content…