← Search

Sean M. Richardson

3 accepted papers

2026

Multiple Streams of Knowledge Retrieval: Enriching and Recalling in Transformers

ICLR 2026poster

When an LLM learns a new fact during finetuning (e.g., new movie releases, updated celebrity gossip, etc.), where does this information go? Are entities enriched with relation information, or do models recall information just-in-time before a prediction? Are ``all of the above'' true with LLMs imple…

Cited by 0SourcecodeScholar
2025

Position: LLM Social Simulations Are a Promising Research Method

ICML 2025poster

Accurate and verifiable large language model (LLM) simulations of human research subjects promise an accessible data source for understanding human behavior and training new AI systems. However, results to date have been limited, and few social scientists have adopted this method. In this position p…

Cited by 4SourcePDFScholar
2025

RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals

ICML 2025poster

Reward models are widely used as proxies for human preferences when aligning or evaluating LLMs. However, reward models are black boxes, and it is often unclear what, exactly, they are actually rewarding. In this paper we develop Rewrite-based Attribute Treatment Estimator (RATE) as an effective met…