← Search

Mehrdad Saberi

5 accepted papers

2026

Revisiting the Past: Data Unlearning with Model State History

ICLR 2026poster

Large language models are trained on massive corpora of web data, which may include private data, copyrighted material, factually inaccurate data, or data that degrades model performance. Eliminating the influence of such problematic datapoints on a model through complete retraining---by repeatedly…

Cited by 0SourcecodeScholar
2025

A Technical Report on “Erasing the Invisible”: The 2024 NeurIPS Competition on Stress Testing Image Watermarks

NeurIPS 2025poster

AI-generated images have become pervasive, raising critical concerns around content authenticity, intellectual property, and the spread of misinformation. Invisible watermarks offer a promising solution for identifying AI-generated images, preserving content provenance without degrading visual quali…

Cited by 0SourceScholar
2025

Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text

NeurIPS 2025poster

The increasing capabilities of Large Language Models (LLMs) have raised concerns about their misuse in AI-generated plagiarism and social engineering. While various AI-generated text detectors have been proposed to mitigate these risks, many remain vulnerable to simple evasion techniques such as par…

Cited by 18SourcecodeScholar
2024

PRIME: Prioritizing Interpretability in Failure Mode Extraction

ICLR 2024poster

In this work, we study the challenge of providing human-understandable descriptions for failure modes in trained image classification models. Existing works address this problem by first identifying clusters (or directions) of incorrectly classified samples in a latent space and then aiming to provi…

Cited by 4SourcePDFScholar
2024

Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks

ICLR 2024poster

In light of recent advancements in generative AI models, it has become essential to distinguish genuine content from AI-generated one to prevent the malicious usage of fake materials as authentic ones and vice versa. Various techniques have been introduced for identifying AI-generated images, with w…