← Search

Conor Houghton

3 accepted papers

2025

Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations

ICML 2025poster

Sparse autoencoders (SAEs) have been successfully used to discover sparse and human-interpretable representations of the latent activations of language models (LLMs). However, we would ultimately like to understand the computations performed by LLMs and not just their representations. The extent to…

Cited by 0SourcePDFScholar
2025

Residual Stream Analysis with Multi-Layer SAEs

ICLR 2025poster

Sparse autoencoders (SAEs) are a promising approach to interpreting the internal representations of transformer language models. However, SAEs are usually trained separately on each transformer layer, making it difficult to use them to study how information flows across layers. To solve this problem…

2019

Adaptive Estimators Show Information Compression in Deep Neural Networks

ICLR 2019poster

To improve how neural networks function it is crucial to understand their learning process. The information bottleneck theory of deep learning proposes that neural networks achieve good generalization by compressing their representations to disregard information that is not relevant to the task. How…

Cited by 52SourcePDFScholar