← Search

Aleksandar Makelov

5 accepted papers

2026

Persona Features Control Emergent Misalignment

ICLR 2026poster

Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. discovered that fine-tuning GPT-4o on intentionally insecure code causes "emergent misalignment," where models give stereotypically mali…

Cited by 0SourcecodeScholar
2025

Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control

ICLR 2025poster

Disentangling model activations into human-interpretable features is a central problem in interpretability. Sparse autoencoders (SAEs) have recently attracted much attention as a scalable unsupervised approach to this problem. However, our imprecise understanding of ground-truth features in realisti…

Cited by 30SourcePDFScholar
2024

Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching

ICLR 2024poster

Mechanistic interpretability aims to attribute high-level model behaviors to specific, interpretable learned features. It is hypothesized that these features manifest as directions or low-dimensional subspaces within activation space. Accordingly, recent studies have explored the identification and…

2023

Rethinking Backdoor Attacks

ICML 2023poster

In a *backdoor attack*, an adversary inserts maliciously constructed backdoor examples into a training set to make the resulting model vulnerable to manipulation. Defending against such attacks involves viewing inserted examples as outliers in the training set and using techniques from robust statis…

Cited by 22SourcePDFScholar
2018

Towards Deep Learning Models Resistant to Adversarial Attacks

ICLR 2018poster

Recent work has demonstrated that neural networks are vulnerable to adversarial examples, i.e., inputs that are almost indistinguishable from natural data and yet classified incorrectly by the network. To address this problem, we study the adversarial robustness of neural networks through the lens o…

Cited by 15225SourcePDFScholar