← Search

Ihsen Alouani

5 accepted papers

2025

Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment

EMNLP 2025

Recent research has shown that carefully crafted jailbreak inputs can induce large language models to produce harmful outputs, despite safety measures such as alignment. It is important to anticipate the range of potential Jailbreak attacks to guide effective defenses and accurate assessment of mode

Cited by 0SourcePDFScholar
2025

Mind the Gap: Detecting Black-box Adversarial Attacks in the Making through Query Update Analysis

CVPR 2025poster

Adversarial attacks remain a significant threat that can jeopardize the integrity of Machine Learning (ML) models. In particular, query-based black-box attacks can generate malicious noise without having access to the victim model's architecture, making them practical in real-world contexts. The com…

2024

DAP: A Dynamic Adversarial Patch for Evading Person Detectors

CVPR 2024poster

Patch-based adversarial attacks were proven to compromise the robustness and reliability of computer vision systems. However their conspicuous and easily detectable nature challenge their practicality in real-world setting. To address this recent work has proposed using Generative Adversarial Networ…

Cited by 30SourcePDFScholar
2024

SSAP: A Shape-Sensitive Adversarial Patch for Comprehensive Disruption of Monocular Depth Estimation in Autonomous Navigation Applications

IROS 2024poster

Monocular depth estimation (MDE) has advanced significantly, primarily through the integration of convolutional neural networks (CNNs) and more recently, Transformers. However, concerns about their susceptibility to adversarial attacks have emerged, especially in safety-critical domains like autonom…

Cited by 7SourceScholar
2023

Jedi: Entropy-Based Localization and Removal of Adversarial Patches

CVPR 2023poster

Real-world adversarial physical patches were recently shown to be successful in compromising state-of-the-art models in a variety of computer vision applications. The most promising defenses that are based on either input gradient or features analyses have been shown to be compromised by recent GAN-…

Cited by 32SourcePDFScholar