2025
What Makes a Good Dataset for Knowledge Distillation?
CVPR 2025poster
Knowledge distillation (KD) has been a popular and effective method for model compression. One important assumption of KD is that the teacher's original dataset will also be available when training the student. However, in situations such as continual learning and distilling large models trained on…