← Search

Wen-Shu Fan

3 accepted papers

2025

Maximizing the Effectiveness of Larger BERT Models for Compression

ACL 2025long

Knowledge distillation (KD) is a widely used approach for BERT compression, where a larger BERT model serves as a teacher to transfer knowledge to a smaller student model. Prior works have found that distilling a larger BERT with superior performance may degrade student’s performance than a smaller…

2024

Revisit the Essence of Distilling Knowledge through Calibration

ICML 2024poster

Knowledge Distillation (KD) has evolved into a practical technology for transferring knowledge from a well-performing model (teacher) to a weak model (student). A counter-intuitive phenomenon known as capacity mismatch has been identified, wherein KD performance may not be good when a better teacher…

Cited by 1SourcePDFScholar
2022

Asymmetric Temperature Scaling Makes Larger Networks Teach Well Again

NeurIPS 2022accept

Knowledge Distillation (KD) aims at transferring the knowledge of a well-performed neural network (the {\it teacher}) to a weaker one (the {\it student}). A peculiar phenomenon is that a more accurate model doesn't necessarily teach better, and temperature adjustment can neither alleviate the mismat…

Cited by 39SourcePDFScholar