← Search

Cheng Hou

2 accepted papers

2024

Dynamic Data Sampler for Cross-Language Transfer Learning in Large Language Models

ICASSP 2024accepted

Large Language Models (LLMs) have gained significant attention in the field of natural language processing (NLP) due to their wide range of applications. However, training LLMs for languages other than English poses significant challenges, due to the difficulty in acquiring large-scale corpus and th…

Cited by 0SourceScholar
2024

Weight-Inherited Distillation for Task-Agnostic BERT Compression

NAACL 2024findings

Knowledge Distillation (KD) is a predominant approach for BERT compression.Previous KD-based methods focus on designing extra alignment losses for the student model to mimic the behavior of the teacher model.These methods transfer the knowledge in an indirect way.In this paper, we propose a novel We…