← Search

Fabio Roli

7 accepted papers

2026

Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization

ICLR 2026poster

Vision–Language Models (VLMs) have become essential for tasks such as image synthesis, captioning, and retrieval by aligning textual and visual information in a shared embedding space. Yet, this flexibility also makes them vulnerable to malicious prompts designed to produce unsafe content, raising c…

Cited by 0SourceScholar
2026

SOM Directions Are Better than One: Multi-Directional Refusal Suppression in Language Models

AAAI 2026technical

Refusal refers to the functional behavior enabling safety-aligned language models to reject harmful or unethical prompts. Following the growing scientific interest in mechanistic interpretability, recent work encoded refusal behavior as a single direction in the model’s latent space; e.g., computed

Cited by 0SourcePDFScholar
2025

AttackBench: Evaluating Gradient-based Attacks for Adversarial Examples

AAAI 2025technical

While novel gradient-based attacks are continuously proposed to improve the optimization of adversarial examples, each is shown to outperform its predecessors using different experimental setups, implementations, and computational budgets, leading to biased and unfair comparisons. In this work, we o…

Cited by 5SourcePDFScholar
2025

TransferBench: Benchmarking Ensemble-based Black-box Transfer Attacks

NeurIPS 2025poster

Ensemble-based black-box transfer attacks optimize adversarial examples on a set of surrogate models, claiming to reach high success rates by querying the (unknown) target model only a few times. In this work, we show that prior evaluations are systematically biased, as such methods are tested only…

Cited by 0SourcecodeScholar
2022

Indicators of Attack Failure: Debugging and Improving Optimization of Adversarial Examples

NeurIPS 2022accept

Evaluating robustness of machine-learning models to adversarial examples is a challenging problem. Many defenses have been shown to provide a false sense of robustness by causing gradient-based attacks to fail, and they have been broken under more rigorous evaluations. Although guidelines and best p…

2021

Fast Minimum-norm Adversarial Attacks through Adaptive Norm Constraints

NeurIPS 2021poster

Evaluating adversarial robustness amounts to finding the minimum perturbation needed to have an input sample misclassified. The inherent complexity of the underlying optimization requires current gradient-based attacks to be carefully tuned, initialized, and possibly executed for many computational…

2015

Is Feature Selection Secure against Training Data Poisoning?

ICML 2015poster

Learning in adversarial settings is becoming an important task for application domains where attackers may inject malicious data into the training set to subvert normal operation of data-driven technologies. Feature selection has been widely used in machine learning for security applications to impr…

Cited by 545SourcePDFScholar