← Search

Edward Stevinson

3 accepted papers

2026

Adversarial Vulnerability from Interference Between Features in Superposition

ICML 2026poster

Why do adversarial examples exist, and why do they transfer between models? Existing explanations appeal to high-dimensional geometry, non-robust patterns in the input, and decision boundary structure, but none provides a representation-level mechanism that explains why specific perturbations succee…

Cited by 0SourceScholar
2026

ContextBench: Modifying Contexts for Targeted Latent Activation and Behaviour Elicitation

ICLR 2026poster

Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of generating targeted, linguistically fluent inputs that activate specific latent features or elicit model behaviours. We…

Cited by 0SourcecodeScholar
2026

Correlations in the Data Lead to Semantically Rich Feature Geometry Under Superposition

ICLR 2026poster

Recent advances in mechanistic interpretability have shown that many features represented by deep learning models can be captured by dictionary learning approaches such as sparse autoencoders. However, our understanding of the structures formed by these internal representations is still limited. Ini…

Cited by 0SourcecodeScholar