← Search

Raffaele Mura

2 accepted papers

2026

SOM Directions Are Better than One: Multi-Directional Refusal Suppression in Language Models

AAAI 2026technical

Refusal refers to the functional behavior enabling safety-aligned language models to reject harmful or unethical prompts. Following the growing scientific interest in mechanistic interpretability, recent work encoded refusal behavior as a single direction in the model’s latent space; e.g., computed

Cited by 0SourcePDFScholar
2025

TransferBench: Benchmarking Ensemble-based Black-box Transfer Attacks

NeurIPS 2025poster

Ensemble-based black-box transfer attacks optimize adversarial examples on a set of surrogate models, claiming to reach high success rates by querying the (unknown) target model only a few times. In this work, we show that prior evaluations are systematically biased, as such methods are tested only…

Cited by 0SourcecodeScholar