2022
Asymmetric Temperature Scaling Makes Larger Networks Teach Well Again
NeurIPS 2022accept
Knowledge Distillation (KD) aims at transferring the knowledge of a well-performed neural network (the {\it teacher}) to a weaker one (the {\it student}). A peculiar phenomenon is that a more accurate model doesn't necessarily teach better, and temperature adjustment can neither alleviate the mismat…