← Search

Yuanchun Wang

2 accepted papers

2023

Are Intermediate Layers and Labels Really Necessary? A General Language Model Distillation Method

ACL 2023findings

The large scale of pre-trained language models poses a challenge for their deployment on various devices, with a growing emphasis on methods to compress these models, particularly knowledge distillation. However, current knowledge distillation methods rely on the model’s intermediate layer features…

2023

GKD: A General Knowledge Distillation Framework for Large-scale Pre-trained Language Model

ACL 2023industry

Currently, the reduction in the parameter scale of large-scale pre-trained language models (PLMs) through knowledge distillation has greatly facilitated their widespread deployment on various devices. However, the deployment of knowledge distillation systems faces great challenges in real-world indu…