← Search

Huadi Zheng

8 accepted papers

2026

Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction Attacks

AAAI 2026technical

Machine learning models constitute valuable intellectual property, yet remain vulnerable to model extraction attacks (MEA), where adversaries replicate their functionality through black-box queries. Model watermarking counters MEAs by embedding forensic markers for ownership verification. Current bl

Cited by 0SourcePDFScholar
2026

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling

ICML 2026poster

Speculative inference accelerates large language model (LLM) decoding but provides no inherent safety guarantees. Existing safety defenses are largely incompatible with speculative inference: they either introduce additional computation or disrupt the draft–verify mechanism, negating acceleration be…

Cited by 0SourceScholar
2026

Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs

ICML 2026poster

Unlearning in large language models (LLMs) aims to remove specified data, but its efficacy is typically assessed with task-level metrics like accuracy and perplexity. We demonstrate that these metrics are often misleading, as models can appear to forget while their original behavior is easily restor…

Cited by 0SourcecodeScholar
2025

A Sample-Level Evaluation and Generative Framework for Model Inversion Attacks

AAAI 2025technical

Model Inversion (MI) attacks, which reconstruct the training dataset of neural networks, pose significant privacy concerns in machine learning. Recent MI attacks have managed to reconstruct realistic label-level private data, such as the general appearance of a target person from all training images…

2025

Multi-Turn Jailbreaking Large Language Models via Attention Shifting

AAAI 2025technical

Large Language Models (LLMs) have achieved significant performance in various natural language processing tasks but also pose safety and ethical threats, thus requiring red teaming and alignment processes to bolster their safety. To effectively exploit these aligned LLMs, recent studies have introdu…

Cited by 0SourcePDFScholar
2025

Reminiscence Attack on Residuals: Exploiting Approximate Machine Unlearning for Privacy

ICCV 2025poster

Machine unlearning enables the removal of specific data from ML models to uphold the *right to be forgotten*. While approximate unlearning algorithms offer efficient alternatives to full retraining, this work reveals that they fail to adequately protect the privacy of unlearned data. In particular,…

Cited by 0SourcePDFScholar
2025

SilentStriker: Toward Stealthy Bit-Flip Attacks on Large Language Models

NeurIPS 2025poster

The rapid adoption of large language models (LLMs) in critical domains has spurred extensive research into their security issues. While input manipulation attacks (e.g., prompt injection) have been well-studied, Bit-Flip Attacks (BFAs)—which exploit hardware vulnerabilities to corrupt model paramete…

Cited by 0SourceScholar
2022

MExMI: Pool-based Active Model Extraction Crossover Membership Inference

NeurIPS 2022accept

With increasing popularity of Machine Learning as a Service (MLaaS), ML models trained from public and proprietary data are deployed in the cloud and deliver prediction services to users. However, as the prediction API becomes a new attack surface, growing concerns have arisen on the confidentiality…

Cited by 13SourcePDFScholar