← Search

Damjan Kalajdzievski

2 accepted papers

2025

The Logical Implication Steering Method for Conditional Interventions on Transformer Generation

ICML 2025poster

The field of mechanistic interpretability in pre-trained transformer models has demonstrated substantial evidence supporting the ''linear representation hypothesis'', which is the idea that high level concepts are encoded as vectors in the space of activations of a model. Studies also show that mode…

Cited by 0SourcePDFScholar
2021

Learning to live with Dale's principle: ANNs with separate excitatory and inhibitory units

ICLR 2021poster

The units in artificial neural networks (ANNs) can be thought of as abstractions of biological neurons, and ANNs are increasingly used in neuroscience research. However, there are many important differences between ANN units and real neurons. One of the most notable is the absence of Dale's principl…

Cited by 47SourcePDFScholar