← Search

Jannes Elstner

1 accepted papers

2025

The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence

ICML 2025poster

The safety alignment of large language models (LLMs) can be circumvented through adversarially crafted inputs, yet the mechanisms by which these attacks bypass safety barriers remain poorly understood. Prior work suggests that a *single* refusal direction in the model's activation space determines w…

Cited by 0SourcePDFScholar