← Search

Giorgio Piras

1 accepted papers

2026

SOM Directions Are Better than One: Multi-Directional Refusal Suppression in Language Models

AAAI 2026technical

Refusal refers to the functional behavior enabling safety-aligned language models to reject harmful or unethical prompts. Following the growing scientific interest in mechanistic interpretability, recent work encoded refusal behavior as a single direction in the model’s latent space; e.g., computed

Cited by 0SourcePDFScholar