← Search

Eric Bigelow

4 accepted papers

2026

Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering

ICML 2026poster

Large language models (LLMs) can be controlled at inference time through prompts (in-context learning) and internal activations (activation steering). Different accounts have been proposed to explain these methods, yet their common goal of controlling model behavior raises the question of whether th…

Cited by 0SourceScholar
2026

Disentangling a Large Language Model’s Computation from its Chain-of-Thought

ICML 2026poster

Do the chains of thought (CoT) of reasoning Large Language Models (LLMs) reflect their internal computation? In this paper, we provide evidence of \textit{performative} CoT, where a model becomes strongly confident in its final answer, but continues generating excess tokens without revealing its int…

Cited by 0SourceScholar
2026

Emergence of Hierarchical Emotion Organization in Large Language Models

ICML 2026poster

As large language models (LLMs) increasingly power conversational agents, understanding how they model users' emotional states is critical for ethical deployment. Inspired by emotion wheels, i.e., a psychological framework that argues emotions organize hierarchically, we analyze probabilistic depend…

Cited by 0SourceScholar
2026

Priors in time: Missing inductive biases for language model interpretability

ICLR 2026poster

A central aim of interpretability tools applied to language models is to recover meaningful concepts from model activations. Existing feature extraction methods focus on single activations regardless of the context, implicitly assuming independence (and therefore stationarity). This leaves open whet…

Cited by 0SourcecodeScholar