2026
Eliciting Harmful Capabilities by Fine-Tuning on Safeguarded Outputs
ICLR 2026poster
Model developers implement safeguards in frontier models to prevent misuse, for example, by employing classifiers to filter dangerous outputs. In this work, we demonstrate that even robustly safeguarded models can be used to elicit harmful capabilities in open-source models through \textit{elicitati…