← Search

Alon Zolfi

7 accepted papers

2025

DIESEL: A Lightweight Inference-Time Safety Enhancement for Language Models

ACL 2025finding

Large language models (LLMs) have demonstrated impressive performance across a wide range of tasks, including open-ended dialogue, driving advancements in virtual assistants and other interactive systems. However, these models often generate outputs misaligned with human values, such as ethical norm…

Cited by 0SourcePDFScholar
2025

Gradient Inversion of Multimodal Models

ICML 2025poster

Federated learning (FL) enables privacy-preserving distributed machine learning by sharing gradients instead of raw data. However, FL remains vulnerable to gradient inversion attacks, in which shared gradients can reveal sensitive training data. Prior research has mainly concentrated on unimodal tas…

Cited by 0SourcePDFScholar
2024

DeSparsify: Adversarial Attack Against Token Sparsification Mechanisms

NeurIPS 2024spotlight

Vision transformers have shown remarkable advancements in the computer vision domain, demonstrating state-of-the-art performance in diverse tasks (e.g., image classification, object detection). However, their high computational requirements grow quadratically with the number of tokens used. Token sp…

Cited by 0SourcePDFScholar
2024

MONTAGE: Monitoring Training for Attribution of Generative Diffusion Models

ECCV 2024poster

"Diffusion models, which revolutionized image generation, are facing challenges related to intellectual property. These challenges arise when a generated image is influenced by copyrighted images from the training data, a plausible scenario in internet-collected data. Hence, pinpointing influential…

2024

Universal Adversarial Attack Against Speaker Recognition Models

ICASSP 2024accepted

In recent years, deep learning-based speaker recognition (SR) models have received a large amount of attention from the machine learning (ML) community. Their increasing popularity derives in large part from their effectiveness in identifying speakers in many security-sensitive applications. Researc…

Cited by 0SourceScholar
2024

YolOOD: Utilizing Object Detection Concepts for Multi-Label Out-of-Distribution Detection

CVPR 2024poster

Out-of-distribution (OOD) detection has attracted a large amount of attention from the machine learning research community in recent years due to its importance in deployed systems. Most of the previous studies focused on the detection of OOD samples in the multi-class classification task. However O…

2021

The Translucent Patch: A Physical and Universal Attack on Object Detectors

CVPR 2021poster

Physical adversarial attacks against object detectors have seen increasing success in recent years. However, these attacks require direct access to the object of interest in order to apply a physical patch. Furthermore, to hide multiple objects, an adversarial patch must be applied to each object. I…

Cited by 129PDFScholar