← Search

Zhong-Yi Lu

3 accepted papers

2024

Over-parameterized Student Model via Tensor Decomposition Boosted Knowledge Distillation

NeurIPS 2024poster

Increased training parameters have enabled large pre-trained models to excel in various downstream tasks. Nevertheless, the extensive computational requirements associated with these models hinder their widespread adoption within the community. We focus on Knowledge Distillation (KD), where a compac…

2022

Parameter-Efficient Mixture-of-Experts Architecture for Pre-trained Language Models

COLING 2022main

Recently, Mixture-of-Experts (short as MoE) architecture has achieved remarkable success in increasing the model capacity of large-scale language models. However, MoE requires incorporating significantly more parameters than the base model being extended. In this paper, we propose building a paramet…

2021

Enabling Lightweight Fine-tuning for Pre-trained Language Model Compression based on Matrix Product Operators

ACL 2021long

This paper presents a novel pre-trained language models (PLM) compression approach based on the matrix product operator (short as MPO) from quantum many-body physics. It can decompose an original matrix into central tensors (containing the core information) and auxiliary tensors (with only a small p…