← Search

Guoqing Jiang

3 accepted papers

2024

How to Trade Off the Quantity and Capacity of Teacher Ensemble: Learning Categorical Distribution to Stochastically Employ a Teacher for Distillation

AAAI 2024technical

We observe two phenomenons with respect to quantity and capacity: 1) more teacher is not always better for multi-teacher knowledge distillation, and 2) stronger teacher is not always better for single-teacher knowledge distillation. To trade off the quantity and capacity of teacher ensemble, in this…

Cited by 1SourcePDFScholar
2023

SKDBERT: Compressing BERT via Stochastic Knowledge Distillation

AAAI 2023technical

In this paper, we propose Stochastic Knowledge Distillation (SKD) to obtain compact BERT-style language model dubbed SKDBERT. In each distillation iteration, SKD samples a teacher model from a pre-defined teacher team, which consists of multiple teacher models with multi-level capacities, to transfe…

Cited by 16SourcePDFScholar
2020

Understanding Why Neural Networks Generalize Well Through GSNR of Parameters

ICLR 2020spotlight

As deep neural networks (DNNs) achieve tremendous success across many application domains, researchers tried to explore in many aspects on why they generalize well. In this paper, we provide a novel perspective on these issues using the gradient signal to noise ratio (GSNR) of parameters during trai…

Cited by 61SourceScholar