← Search

CHING-CHIA KAO

4 accepted papers

2026

Submodular Optimization for Minimal Augmentation in Robust Language Model Alignment

ICML 2026poster

Safety alignment of large language models is fragile: even small fine-tuning perturbations elastically revert behaviors toward those of the pre-training, with degradation inversely proportional to the size of the alignment set. We ask how to achieve safety alignment with \emph{minimal augmentation}.…

Cited by 0SourceScholar
2025

Safety Depth in Large Language Models: A Markov Chain Perspective

NeurIPS 2025poster

Large Language Models (LLMs) are increasingly adopted in high-stakes scenarios, yet their safety mechanisms often remain fragile. Simple jailbreak prompts or even benign fine-tuning can bypass internal safeguards, underscoring the need to understand the failure modes of current safety strategies. R…

Cited by 0SourceScholar
2024

Defending against Clean-Image Backdoor Attack in Multi-Label Classification

ICASSP 2024accepted

Deep neural networks (DNNs) are known to be vulnerable to backdoor attacks. Specifically, the attacker endeavors to implant backdoors in the DNN model by injecting a set of poisoning samples such that the malicious model predicts target labels once the backdoor is triggered. The clean-image attack h…

Cited by 0SourceScholar
2022

DPGEN: Differentially Private Generative Energy-Guided Network for Natural Image Synthesis

CVPR 2022oral

Despite an increased demand for valuable data, the privacy concerns associated with sensitive datasets present a barrier to data sharing. One may use differentially private generative models to generate synthetic data. Unfortunately, generators are typically restricted to generating images of low-re…

Cited by 28PDFcodeScholar