Dataset Distillation for Memorized Data: Soft Labels can Leak Held-Out Teacher Knowledge
Dataset distillation aims to compress training data into fewer examples via a teacher, from which a student can learn effectively. While its success is often attributed to structure in the data, modern neural networks also memorize specific facts, but if and how such memorized information can be tra…