← Search

Joseph Isaac Bloom

4 accepted papers

2026

ContextBench: Modifying Contexts for Targeted Latent Activation and Behaviour Elicitation

ICLR 2026poster

Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of generating targeted, linguistically fluent inputs that activate specific latent features or elicit model behaviours. We…

Cited by 0SourcecodeScholar
2025

A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders

NeurIPS 2025oral

Sparse Autoencoders (SAEs) aim to decompose the activation space of large language models (LLMs) into human-interpretable latent directions or features. As we increase the number of features in the SAE, hierarchical features tend to split into finer features (“math” may split into “algebra”, “geomet…

Cited by 0SourcecodeScholar
2025

SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

ICML 2025poster

Sparse autoencoders (SAEs) are a popular technique for interpreting language model activations, and there is extensive recent work on improving SAE effectiveness. However, most prior work evaluates progress using unsupervised proxy metrics with unclear practical relevance. We introduce SAEBench, a c…

Cited by 0SourcePDFScholar
2025

Sparse Autoencoders Do Not Find Canonical Units of Analysis

ICLR 2025poster

A common goal of mechanistic interpretability is to decompose the activations of neural networks into features: interpretable properties of the input computed by the model. Sparse autoencoders (SAEs) are a popular method for finding these features in LLMs, and it has been postulated that they can be…

Cited by 1SourcePDFScholar