← Search

Achyuta Rajaram

3 accepted papers

2026

Persona Features Control Emergent Misalignment

ICLR 2026poster

Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. discovered that fine-tuning GPT-4o on intentionally insecure code causes "emergent misalignment," where models give stereotypically mali…

Cited by 0SourcecodeScholar
2026

Weight-sparse transformers have interpretable circuits

ICML 2026poster

Finding human-understandable circuits in language models is a central goal of the field of mechanistic interpretability. We train models to have more understandable circuits by constraining most of their weights to be zeros, so that each neuron only has a few connections. To recover fine-grained cir…

Cited by 0SourceScholar
2024

A Multimodal Automated Interpretability Agent

ICML 2024poster

This paper describes MAIA, a Multimodal Automated Interpretability Agent. MAIA is a system that uses neural models to automate neural model understanding tasks like feature interpretation and failure mode discovery. It equips a pre-trained vision-language model with a set of tools that support itera…

Cited by 67SourcePDFScholar