← Search

Michael Lan

2 accepted papers

2025

Activation Space Interventions Can Be Transferred Between Large Language Models

ICML 2025poster

The study of representation universality in AI models reveals growing convergence across domains, modalities, and architectures. However, the practical applications of representation universality remain largely unexplored. We bridge this gap by demonstrating that safety interventions can be transfer…

2024

Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models

EMNLP 2024main

While transformer models exhibit strong capabilities on linguistic tasks, their complex architectures make them difficult to interpret. Recent work has aimed to reverse engineer transformer models into human-readable representations called circuits that implement algorithmic functions. We extend thi…