← Search

Antonio Emanuele Cinà

4 accepted papers

2026

Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization

ICLR 2026poster

Vision–Language Models (VLMs) have become essential for tasks such as image synthesis, captioning, and retrieval by aligning textual and visual information in a shared embedding space. Yet, this flexibility also makes them vulnerable to malicious prompts designed to produce unsafe content, raising c…

Cited by 0SourceScholar
2025

$\sigma$-zero: Gradient-based Optimization of $\ell_0$-norm Adversarial Examples

ICLR 2025poster

Evaluating the adversarial robustness of deep networks to gradient-based attacks is challenging. While most attacks consider $\ell_2$- and $\ell_\infty$-norm constraints to craft input perturbations, only a few investigate sparse $\ell_1$- and $\ell_0$-norm attacks. In particular, $\ell_0$-norm atta…

Cited by 0SourcePDFScholar
2025

AttackBench: Evaluating Gradient-based Attacks for Adversarial Examples

AAAI 2025technical

While novel gradient-based attacks are continuously proposed to improve the optimization of adversarial examples, each is shown to outperform its predecessors using different experimental setups, implementations, and computational budgets, leading to biased and unfair comparisons. In this work, we o…

Cited by 5SourcePDFScholar
2025

TransferBench: Benchmarking Ensemble-based Black-box Transfer Attacks

NeurIPS 2025poster

Ensemble-based black-box transfer attacks optimize adversarial examples on a set of surrogate models, claiming to reach high success rates by querying the (unknown) target model only a few times. In this work, we show that prior evaluations are systematically biased, as such methods are tested only…

Cited by 0SourcecodeScholar