← Search

Chia-Mu Yu

19 accepted papers

2026

Submodular Optimization for Minimal Augmentation in Robust Language Model Alignment

ICML 2026poster

Safety alignment of large language models is fragile: even small fine-tuning perturbations elastically revert behaviors toward those of the pre-training, with degradation inversely proportional to the size of the alignment set. We ask how to achieve safety alignment with \emph{minimal augmentation}.…

Cited by 0SourceScholar
2025

Differentially Private Fine-Tuning of Diffusion Models

ICCV 2025poster

Generative AI models, particularly diffusion models (DMs), have demonstrated exceptional capabilities in high-quality image synthesis. However, their large memorization capacity raises significant privacy concerns, especially when trained on sensitive datasets. This paper introduces DP-LoRA, a surpr…

2025

Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge

EMNLP 2025

Large language models (LLMs) demonstrate strong task-specific capabilities through fine-tuning, but merging multiple fine-tuned models often leads to degraded performance due to overlapping instruction-following components. Task Arithmetic (TA), which combines task vectors derived from fine-tuning,

2025

Safety Depth in Large Language Models: A Markov Chain Perspective

NeurIPS 2025poster

Large Language Models (LLMs) are increasingly adopted in high-stakes scenarios, yet their safety mechanisms often remain fragile. Simple jailbreak prompts or even benign fine-tuning can bypass internal safeguards, underscoring the need to understand the failure modes of current safety strategies. R…

Cited by 0SourceScholar
2025

VP-NTK: Exploring the Benefits of Visual Prompting in Differentially Private Data Synthesis

ICASSP 2025accepted

Differentially private (DP) synthetic data has become the de facto standard for releasing sensitive data. However, many DP generative models suffer from the low utility of synthetic data, especially for high-resolution images. On the other hand, one of the emerging techniques in parameter efficient…

Cited by 0SourceScholar
2024

Defending against Clean-Image Backdoor Attack in Multi-Label Classification

ICASSP 2024accepted

Deep neural networks (DNNs) are known to be vulnerable to backdoor attacks. Specifically, the attacker endeavors to implant backdoors in the DNN model by injecting a set of poisoning samples such that the malicious model predicts target labels once the backdoor is triggered. The clean-image attack h…

Cited by 0SourceScholar
2024

Rethinking Backdoor Attacks on Dataset Distillation: A Kernel Method Perspective

ICLR 2024poster

Dataset distillation offers a potential means to enhance data efficiency in deep learning. Recent studies have shown its ability to counteract backdoor risks present in original training samples. In this study, we delve into the theoretical aspects of backdoor attacks and dataset distillation based…

Cited by 10SourcePDFScholar
2024

Ring-A-Bell! How Reliable are Concept Removal Methods For Diffusion Models?

ICLR 2024poster

Diffusion models for text-to-image (T2I) synthesis, such as Stable Diffusion (SD), have recently demonstrated exceptional capabilities for generating high-quality content. However, this progress has raised several concerns of potential misuse, particularly in creating copyrighted, prohibited, and re…

2024

Safe LoRA: The Silver Lining of Reducing Safety Risks when Finetuning Large Language Models

NeurIPS 2024poster

While large language models (LLMs) such as Llama-2 or GPT-4 have shown impressive zero-shot performance, fine-tuning is still necessary to enhance their performance for customized datasets, domain-specific tasks, or other private needs. However, fine-tuning all parameters of LLMs requires significan…

2023

Certified Robustness of Quantum Classifiers Against Adversarial Examples Through Quantum Noise

ICASSP 2023accepted

Recently, quantum classifiers have been known to be vulnerable to adversarial attacks, where quantum classifiers are fooled by imperceptible noises to have misclassification. In this paper, we propose one first theoretical study that utilizing the added quantum random rotation noise can improve the…

Cited by 0SourceScholar
2023

Exploring the Benefits of Visual Prompting in Differential Privacy

ICCV 2023poster

Visual Prompting (VP) is an emerging and powerful technique that allows sample-efficient adaptation to downstream tasks by engineering a well-trained frozen source model. In this work, we explore the benefits of VP in constructing compelling neural network classifiers with differential privacy (DP).…

Cited by 19PDFcodeScholar
2022

Adversarial Examples Can Be Effective Data Augmentation for Unsupervised Machine Learning

AAAI 2022technical

Adversarial examples causing evasive predictions are widely used to evaluate and improve the robustness of machine learning models. However, current studies focus on supervised learning tasks, relying on the ground truth data label, a targeted objective, or supervision from a trained classifier. In…

2022

DPGEN: Differentially Private Generative Energy-Guided Network for Natural Image Synthesis

CVPR 2022oral

Despite an increased demand for valuable data, the privacy concerns associated with sensitive datasets present a barrier to data sharing. One may use differentially private generative models to generate synthetic data. Unfortunately, generators are typically restricted to generating images of low-re…

Cited by 28PDFcodeScholar
2021

CAFE: Catastrophic Data Leakage in Vertical Federated Learning

NeurIPS 2021poster

Recent studies show that private training data can be leaked through the gradients sharing mechanism deployed in distributed machine learning systems, such as federated learning (FL). Increasing batch size to complicate data recovery is often viewed as a promising defense strategy against data leaka…

2021

Formalizing Generalization and Adversarial Robustness of Neural Networks to Weight Perturbations

NeurIPS 2021poster

Studying the sensitivity of weight perturbation in neural networks and its impacts on model performance, including generalization and robustness, is an active research topic due to its implications on a wide range of machine learning tasks such as model compression, generalization gap assessment, an…

Cited by 29SourcePDFScholar
2021

Perceptual Indistinguishability-Net (PI-Net): Facial Image Obfuscation With Manipulable Semantics

CVPR 2021poster

With the growing use of camera devices, the industry has many image datasets that provide more opportunities for collaboration between the machine learning community and industry. However, the sensitive information in the datasets discourages data owners from releasing these datasets. Despite recent…

Cited by 51PDFcodeScholar