SOM Directions Are Better than One: Multi-Directional Refusal Suppression in Language Models
Giorgio Piras, Raffaele Mura, Fabio Brau, Luca Oneto, Fabio Roli, Battista Biggio
Abstract
Refusal refers to the functional behavior enabling safety-aligned language models to reject harmful or unethical prompts. Following the growing scientific interest in mechanistic interpretability, recent work encoded refusal behavior as a single direction in the model’s latent space; e.g., computed as the difference between the centroids of harmful and harmless prompt representations. However, emerging evidence suggests that concepts in LLMs often appear to be encoded as a low-dimensional manifold embedded in the high-dimensional latent space. Motivated by these findings, we propose a novel method leveraging Self-Organizing Maps (SOMs) to extract multiple refusal directions. To this end, we first prove that SOMs generalize the prior work
BibTeX
@inproceedings{aaai2026_somdirectionsare,
title = {SOM Directions Are Better than One: Multi-Directional Refusal Suppression in Language Models},
author = {Giorgio Piras and Raffaele Mura and Fabio Brau and Luca Oneto and Fabio Roli and Battista Biggio},
booktitle = {AAAI 2026},
year = {2026}
}