Multi-stage Distillation Framework for Cross-Lingual Semantic Similarity Matching
Kunbo Ding, Weijie Liu, Yuejian Fang, Zhe Zhao, Qi Ju, Xuefeng Yang, Rong Tian, Zhu Tao
Abstract
Previous studies have proved that cross-lingual knowledge distillation can significantly improve the performance of pre-trained models for cross-lingual similarity matching tasks. However, the student model needs to be large in this operation. Otherwise, its performance will drop sharply, thus making it impractical to be deployed to memory-limited devices. To address this issue, we delve into cross-lingual knowledge distillation and propose a multi-stage distillation framework for constructing a small-size but high-performance cross-lingual model. In our framework, contrastive learning, bottleneck, and parameter recurrent strategies are delicately combined to prevent performance from being compromised during the compression process. The experimental results demonstrate that our method can compress the size of XLM-R and MiniLM by more than 50%, while the performance is only reduced by about 1%.
BibTeX
@inproceedings{ding-etal-2022-multi,
title = "Multi-stage Distillation Framework for Cross-Lingual Semantic Similarity Matching",
author = "Ding, Kunbo and
Liu, Weijie and
Fang, Yuejian and
Zhao, Zhe and
Ju, Qi and
Yang, Xuefeng and
Tian, Rong and
Tao, Zhu and
Liu, Haoyan and
Guo, Han and
Bai, Xingyu and
Mao, Weiquan and
Li, Yudong and
Guo, Weigang and
Wu, Taiqiang and
Sun, Ningyuan",
editor = "Carpuat, Marine and
de Marneffe, Marie-Catherine and
Meza Ruiz, Ivan Vladimir",
booktitle = "Findings of the Association for Computational Linguistics: NAACL 2022",
month = jul,
year = "2022",
address = "Seattle, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.findings-naacl.167/",
doi = "10.18653/v1/2022.findings-naacl.167",
pages = "2171--2181"
}