← Search

Sarah Ball

4 accepted papers

2026

On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment

ICLR 2026poster

With the increased deployment of large language models (LLMs), one concern is their potential misuse for generating harmful content. Our work studies the alignment challenge, with a focus on filters to prevent the generation of unsafe information. Two natural points of intervention are the filtering…

Cited by 0SourcecodeScholar
2026

Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting

ICML 2026poster

Large language models are increasingly used to predict human preferences in both scientific and business endeavors, yet current approaches rely exclusively on analyzing model outputs without considering the underlying mechanisms. Using election forecasting as a test case, we introduce *mechanistic f…

Cited by 0SourceScholar