2022
Teach Less, Learn More: On the Undistillable Classes in Knowledge Distillation
NeurIPS 2022accept
Knowledge distillation (KD) can effectively compress neural networks by training a smaller network (student) to simulate the behavior of a larger one (teacher). A counter-intuitive observation is that a more expansive teacher does not make a better student, but the reasons for this phenomenon remain…