← Search

Jackson Kaunismaa

1 accepted papers

2026

Eliciting Harmful Capabilities by Fine-Tuning on Safeguarded Outputs

ICLR 2026poster

Model developers implement safeguards in frontier models to prevent misuse, for example, by employing classifiers to filter dangerous outputs. In this work, we demonstrate that even robustly safeguarded models can be used to elicit harmful capabilities in open-source models through \textit{elicitati…

Cited by 0SourceScholar