← Search

Shaoming Song

2 accepted papers

2022

Asymmetric Temperature Scaling Makes Larger Networks Teach Well Again

NeurIPS 2022accept

Knowledge Distillation (KD) aims at transferring the knowledge of a well-performed neural network (the {\it teacher}) to a weaker one (the {\it student}). A peculiar phenomenon is that a more accurate model doesn't necessarily teach better, and temperature adjustment can neither alleviate the mismat…

Cited by 39SourcePDFScholar
2022

Federated Learning With Position-Aware Neurons

CVPR 2022poster

Federated Learning (FL) fuses collaborative models from local nodes without centralizing users' data. The permutation invariance property of neural networks and the non-i.i.d. data across clients make the locally updated parameters imprecisely aligned, disabling the coordinate-based parameter averag…

Cited by 44PDFcodeScholar