← Search

Sonia Krishna Murthy

4 accepted papers

2026

Priors in time: Missing inductive biases for language model interpretability

ICLR 2026poster

A central aim of interpretability tools applied to language models is to recover meaningful concepts from model activations. Existing feature extraction methods focus on single activations regardless of the context, implicitly assuming independence (and therefore stationarity). This leaves open whet…

Cited by 0SourcecodeScholar
2026

Using cognitive models to reveal value trade-offs in language models

ICLR 2026poster

Value trade-offs are an integral part of human decision-making and language use, however, current tools for interpreting such dynamic and multi-faceted notions of values in LLMs are limited. In cognitive science, so-called “cognitive models” provide formal accounts of such trade-offs in humans, by m…

Cited by 0SourcecodeScholar
2025

One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity

NAACL 2025long

Researchers in social science and psychology have recently proposed using large language models (LLMs) as replacements for humans in behavioral research. In addition to arguments about whether LLMs accurately capture population-level patterns, this has raised questions about whether LLMs capture hum…

2023

Comparing the Evaluation and Production of Loophole Behavior in Humans and Large Language Models

EMNLP 2023long findings

In law, lore, and everyday life, loopholes are commonplace. When people exploit a loophole, they understand the intended meaning or goal of another person, but choose to go with a different interpretation. Past and current AI research has shown that artificial intelligence engages in what seems supe…

Cited by 0SourceScholar