2026
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
ICLR 2026poster
Language models can transfer hidden biases during distillation. For example, a teacher that "likes owls" can make its student "like owls" too, even when the training data consists only of lists of numbers. This surprising phenomenon is called *subliminal learning*. Subliminal learning can be expecte…