← Search

Nicholas Goldowsky-Dill

2 accepted papers

2025

Detecting Strategic Deception with Linear Probes

ICML 2025poster

AI models might use deceptive strategies as part of scheming or misaligned behaviour. Monitoring outputs alone is insufficient, since the AI might produce seemingly benign outputs while its internal reasoning is misaligned. We thus evaluate if linear probes can robustly detect deception by monitorin…

2024

Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning

NeurIPS 2024poster

Identifying the features learned by neural networks is a core challenge in mechanistic interpretability. Sparse autoencoders (SAEs), which learn a sparse, overcomplete dictionary that reconstructs a network's internal activations, have been used to identify these features. However, SAEs may learn mo…