Submodular Optimization for Minimal Augmentation in Robust Language Model Alignment
Safety alignment of large language models is fragile: even small fine-tuning perturbations elastically revert behaviors toward those of the pre-training, with degradation inversely proportional to the size of the alignment set. We ask how to achieve safety alignment with \emph{minimal augmentation}.…