← Search

Lucy Farnik

4 accepted papers

2026

Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers

ICLR 2026poster

Sparse autoencoders (SAEs) are widely used to extract sparse, interpretable latents from transformer activations. We test whether commonly used SAE quality metrics and automatic explanation pipelines can distinguish trained transformers from randomly initialized ones (e.g., where parameters are samp…

Cited by 0SourceScholar
2025

Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations

ICML 2025poster

Sparse autoencoders (SAEs) have been successfully used to discover sparse and human-interpretable representations of the latent activations of language models (LLMs). However, we would ultimately like to understand the computations performed by LLMs and not just their representations. The extent to…

Cited by 0SourcePDFScholar
2025

Residual Stream Analysis with Multi-Layer SAEs

ICLR 2025poster

Sparse autoencoders (SAEs) are a promising approach to interpreting the internal representations of transformer language models. However, SAEs are usually trained separately on each transformer layer, making it difficult to use them to study how information flows across layers. To solve this problem…

2024

STARC: A General Framework For Quantifying Differences Between Reward Functions

ICLR 2024poster

In order to solve a task using reinforcement learning, it is necessary to first formalise the goal of that task as a *reward function*. However, for many real-world tasks, it is very difficult to manually specify a reward function that never incentivises undesirable behaviour. As a result, it is inc…

Cited by 9SourcePDFScholar