← Search

Daniel Wurgaft

3 accepted papers

2026

Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering

ICML 2026poster

Large language models (LLMs) can be controlled at inference time through prompts (in-context learning) and internal activations (activation steering). Different accounts have been proposed to explain these methods, yet their common goal of controlling model behavior raises the question of whether th…

Cited by 0SourceScholar
2026

Priors in time: Missing inductive biases for language model interpretability

ICLR 2026poster

A central aim of interpretability tools applied to language models is to recover meaningful concepts from model activations. Existing feature extraction methods focus on single activations regardless of the context, implicitly assuming independence (and therefore stationarity). This leaves open whet…

Cited by 0SourcecodeScholar
2025

In-Context Learning Strategies Emerge Rationally

NeurIPS 2025poster

Recent work analyzing in-context learning (ICL) has identified a broad set of strategies that describe model behavior in different experimental conditions. We aim to unify these findings by asking why a model learns these disparate strategies in the first place. Specifically, we start with the obser…

Cited by 0SourceScholar