← Search

Fabio Brau

5 accepted papers

2026

SOM Directions Are Better than One: Multi-Directional Refusal Suppression in Language Models

AAAI 2026technical

Refusal refers to the functional behavior enabling safety-aligned language models to reject harmful or unethical prompts. Following the growing scientific interest in mechanistic interpretability, recent work encoded refusal behavior as a single direction in the model’s latent space; e.g., computed

Cited by 0SourcePDFScholar
2025

TransferBench: Benchmarking Ensemble-based Black-box Transfer Attacks

NeurIPS 2025poster

Ensemble-based black-box transfer attacks optimize adversarial examples on a set of surrogate models, claiming to reach high success rates by querying the (unknown) target model only a few times. In this work, we show that prior evaluations are systematically biased, as such methods are tested only…

Cited by 0SourcecodeScholar
2024

1-Lipschitz Layers Compared: Memory Speed and Certifiable Robustness

CVPR 2024poster

The robustness of neural networks against input perturbations with bounded magnitude represents a serious concern in the deployment of deep learning models in safety-critical systems. Recently the scientific community has focused on enhancing certifiable robustness guarantees by crafting \ols neural…

2023

Defending from Physically-Realizable Adversarial Attacks through Internal Over-Activation Analysis

AAAI 2023technical

This work presents Z-Mask, an effective and deterministic strategy to improve the adversarial robustness of convolutional networks against physically-realizable adversarial attacks. The presented defense relies on specific Z-score analysis performed on the internal network features to detect and mas…

Cited by 14SourcePDFScholar
2023

Robust-by-Design Classification via Unitary-Gradient Neural Networks

AAAI 2023technical

The use of neural networks in safety-critical systems requires safe and robust models, due to the existence of adversarial attacks. Knowing the minimal adversarial perturbation of any input x, or, equivalently, knowing the distance of x from the classification boundary, allows evaluating the classif…

Cited by 6SourcePDFScholar