2026
Extracting alignment data in open models
Federico Barbero, Xiangming Gu, Christopher A. Choquette Choo, Chawin Sitawarin, Matthew Jagielski, Itay Yona +3
ICML 2026poster
In this work, we show that it is possible to extract significant amounts of alignment training data from a post-trained model -- useful to steer the model to improve certain capabilities such as long-context reasoning, safety, instruction following, and maths. While the majority of related work on m…