← Search

Anna Soligo

2 accepted papers

2026

Emergent Misalignment is Easy, Narrow Misalignment is Hard

ICLR 2026poster

Finetuning large language models on narrowly harmful datasets can cause them to become emergently misaligned, giving stereotypically `evil' responses across diverse unrelated settings. Concerningly, a pre-registered survey of experts failed to predict this result, highlighting our poor understanding…

Cited by 0SourcecodeScholar
2025

Inducing, Detecting and Characterising Neural Modules: A Pipeline for Functional Interpretability in Reinforcement Learning

ICML 2025poster

Interpretability is crucial for ensuring RL systems align with human values. However, it remains challenging to achieve in complex decision making domains. Existing methods frequently attempt interpretability at the level of fundamental model units, such as neurons or decision nodes: an approach whi…

Cited by 0SourcePDFScholar