← Search

Mazda Moayeri

11 accepted papers

2025

DyePack: Provably Flagging Test Set Contamination in LLMs Using Backdoors

EMNLP 2025

Open benchmarks are essential for evaluating and advancing large language models, offering reproducibility and transparency. However, their accessibility makes them likely targets of test set contamination. In this work, we introduce **DyePack**, a framework that leverages backdoor attacks to identi

2025

Rethinking Artistic Copyright Infringements In the Era Of Text-to-Image Generative Models

ICLR 2025poster

The advent of text-to-image generative models has led artists to worry that their individual styles may be copied, creating a pressing need to reconsider the lack of protection for artistic styles under copyright law. This requires answering challenging questions, like what defines style and what co…

Cited by 4SourcePDFScholar
2025

Unearthing Skill-level Insights for Understanding Trade-offs of Foundation Models

ICLR 2025poster

With models getting stronger, evaluations have grown more complex, testing multiple skills in one benchmark and even in the same instance at once. However, skill-wise performance is obscured when inspecting aggregate accuracy, under-utilizing the rich signal modern benchmarks contain. We propose an…

Cited by 2SourcePDFScholar
2024

PRIME: Prioritizing Interpretability in Failure Mode Extraction

ICLR 2024poster

In this work, we study the challenge of providing human-understandable descriptions for failure modes in trained image classification models. Existing works address this problem by first identifying clusters (or directions) of incorrectly classified samples in a latent space and then aiming to provi…

Cited by 4SourcePDFScholar
2023

A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance Estimation

NeurIPS 2023spotlight

In recent years, concept-based approaches have emerged as some of the most promising explainability methods to help us interpret the decisions of Artificial Neural Networks (ANNs). These methods seek to discover intelligible visual ``concepts'' buried within the complex patterns of ANN activations i…

Cited by 56SourcePDFScholar
2023

Spuriosity Rankings: Sorting Data to Measure and Mitigate Biases

NeurIPS 2023spotlight

We present a simple but effective method to measure and mitigate model biases caused by reliance on spurious cues. Instead of requiring costly changes to one's data or model training, our method better utilizes the data one already has by sorting them. Specifically, we rank images within their class…

Cited by 15SourcePDFScholar
2022

A Comprehensive Study of Image Classification Model Sensitivity to Foregrounds, Backgrounds, and Visual Attributes

CVPR 2022oral

While datasets with single-label supervision have propelled rapid advances in image classification, additional annotations are necessary in order to quantitatively assess how models make predictions. To this end, for a subset of ImageNet samples, we collect segmentation masks for the entire object a…

Cited by 63PDFcodeScholar
2022

Explicit Tradeoffs between Adversarial and Natural Distributional Robustness

NeurIPS 2022accept

Several existing works study either adversarial or natural distributional robustness of deep neural networks separately. In practice, however, models need to enjoy both types of robustness to ensure reliability. In this work, we bridge this gap and show that in fact, {\it explicit tradeoffs} exist b…

Cited by 25SourcePDFScholar
2021

Sample Efficient Detection and Classification of Adversarial Attacks via Self-Supervised Embeddings

ICCV 2021poster

Adversarial robustness of deep models is pivotal in ensuring safe deployment in real world settings, but most modern defenses have narrow scope and expensive costs. In this paper, we propose a self-supervised method to detect adversarial attacks and classify them to their respective threat models, b…

Cited by 31PDFScholar