← Search

Dhanya Sridhar

8 accepted papers

2026

Position: Causality is Key for Interpretability Claims to Generalise

ICML 2026poster

Interpretability research on large language models (LLMs) has produced methods that align model components to high-level concepts, yet their use has been accompanied by recurring failures: findings that do not generalise, and causal language that outruns the evidence. Our position is that Pearl’s ca…

Cited by 0SourceScholar
2025

Causal Differentiating Concepts: Interpreting LM Behavior via Causal Representation Learning

NeurIPS 2025spotlight

Language model activations entangle concepts that mediate their behavior, making it difficult to interpret these factors, which has implications for generalizability and robustness. We introduce an approach for disentangling these concepts without supervision. Existing methods for concept discovery…

Cited by 0SourceScholar
2025

Does learning the right latent variables necessarily improve in-context learning?

ICML 2025poster

Large autoregressive models like Transformers can solve tasks through in-context learning (ICL) without learning new weights, suggesting avenues for efficiently solving new tasks. For many tasks, e.g., linear regression, the data factorizes: examples are independent given a task latent that generate…

2025

In-Context Learning and Occam's Razor

ICML 2025poster

A central goal of machine learning is generalization. While the No Free Lunch Theorem states that we cannot obtain theoretical guarantees for generalization without further assumptions, in practice we observe that simple models which explain the training data generalize best—a principle called Occam…

2022

On the Assumptions of Synthetic Control Methods

AISTATS 2022poster

Synthetic control (SC) methods have been widely applied to estimate the causal effect of large-scale interventions, e.g., the state-wide effect of a change in policy. The idea of synthetic controls is to approximate one unit’s counterfactual outcomes using a weighted combination of some other units’…

2021

Causal Effects of Linguistic Properties

NAACL 2021long

We consider the problem of using observational data to estimate the causal effects of linguistic properties. For example, does writing a complaint politely lead to a faster response time? How much will a positive product review increase sales? This paper addresses two technical challenges related to…

2021

Valid Causal Inference with (Some) Invalid Instruments

ICML 2021spotlight

Instrumental variable methods provide a powerful approach to estimating causal effects in the presence of unobserved confounding. But a key challenge when applying them is the reliance on untestable "exclusion" assumptions that rule out any relationship between the instrument variable and the respon…

Cited by 30SourcePDFScholar