← Search

Ali Shahin Shamsabadi

14 accepted papers

2025

Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention

ICML 2025poster

Cautious predictions—where a machine learning model abstains when uncertain—are crucial for limiting harmful errors in safety-critical applications. In this work, we identify a novel threat: a dishonest institution can exploit these mechanisms to discriminate or unjustly deny services under the guis…

2025

Context-Aware Membership Inference Attacks against Pre-trained Large Language Models

EMNLP 2025

Membership Inference Attacks (MIAs) on pre-trained Large Language Models (LLMs) aim at determining if a data point was part of the model’s training set. Prior MIAs that are built for classification models fail at LLMs, due to ignoring the generative nature of LLMs across token sequences. In this pap

Cited by 0SourcePDFScholar
2025

Membership and Memorization in LLM Knowledge Distillation

EMNLP 2025

Recent advances in Knowledge Distillation (KD) aim to mitigate the high computational demands of Large Language Models (LLMs) by transferring knowledge from a large ”teacher” to a smaller ”student” model. However, students may inherit the teacher’s privacy when the teacher is trained on private data

Cited by 0SourcePDFScholar
2025

Pin the Tail on the Model: Blindfolded Repair of User-Flagged Failures in Text-to-Image Services

NeurIPS 2025poster

Diffusion models are increasingly deployed in real-world text-to-image services. These models, however, encode implicit assumptions about the world based on web-scraped image-caption pairs used during training. Over time, such assumptions may become outdated, incorrect, or socially biased--leading t…

Cited by 0SourceScholar
2025

Secure and Confidential Certificates of Online Fairness

NeurIPS 2025poster

The "black-box service model" enables ML service providers to serve clients while keeping their intellectual property and client data confidential. Confidentiality is critical for delivering ML services legally and responsibly, but makes it difficult for outside parties to verify important model pro…

Cited by 0SourceScholar
2024

Confidential-DPproof: Confidential Proof of Differentially Private Training

ICLR 2024spotlight

Post hoc privacy auditing techniques can be used to test the privacy guarantees of a model, but come with several limitations: (i) they can only establish lower bounds on the privacy loss, (ii) the intermediate model updates and some data must be shared with the auditor to get a better approximation…

Cited by 5SourcePDFScholar
2023

Confidential-PROFITT: Confidential PROof of FaIr Training of Trees

ICLR 2023top-5%

Post hoc auditing of model fairness suffers from potential drawbacks: (1) auditing may be highly sensitive to the test samples chosen; (2) the model and/or its training data may need to be shared with an auditor thereby breaking confidentiality. We address these issues by instead providing a certifi…

Cited by 22SourcePDFScholar
2023

Mnemonist: Locating Model Parameters that Memorize Training Examples

UAI 2023poster

Recent work has shown that an adversary can reconstruct training examples given access to the parameters of a deep learning image classification model. We show that the quality of reconstruction depends heavily on the type of activation functions used. In particular, we show that ReLU activations le…

Cited by 2SourcePDFScholar
2022

A Zest of LIME: Towards Architecture-Independent Model Distances

ICLR 2022poster

Definitions of the distance between two machine learning models either characterize the similarity of the models' predictions or of their weights. While similarity of weights is attractive because it implies similarity of predictions in the limit, it suffers from being inapplicable to comparing mode…

Cited by 26SourcePDFScholar
2022

Washing The Unwashable : On The (Im)possibility of Fairwashing Detection

NeurIPS 2022accept

The use of black-box models (e.g., deep neural networks) in high-stakes decision-making systems, whose internal logic is complex, raises the need for providing explanations about their decisions. Model explanation techniques mitigate this problem by generating an interpretable and high-fidelity surr…

2021

FoolHD: Fooling Speaker Identification by Highly Imperceptible Adversarial Disturbances

ICASSP 2021accepted

Speaker identification models are vulnerable to carefully designed adversarial perturbations of their input signals that induce misclassification. In this work, we propose a white-box steganography-inspired adversarial attack that generates imperceptible adversarial perturbations against a speaker i…

Cited by 0SourceScholar
2019

Scene Privacy Protection

ICASSP 2019accepted

Images shared on social media are routinely analysed by classifiers for content annotation and user profiling. These automatic inferences reveal to the service provider sensitive information that a naive user might want to keep private. To address this problem, we present a method designed to distor…

Cited by 0SourceScholar