← Search

Siyue Wu

3 accepted papers

2023

AD-KD: Attribution-Driven Knowledge Distillation for Language Model Compression

ACL 2023long

Knowledge distillation has attracted a great deal of interest recently to compress large language models. However, existing knowledge distillation methods suffer from two limitations. First, the student model simply imitates the teacher’s behavior while ignoring the reasoning behind it. Second, thes…

2023

MCC-KD: Multi-CoT Consistent Knowledge Distillation

EMNLP 2023long findings

Large language models (LLMs) have showcased remarkable capabilities in complex reasoning through chain of thought (CoT) prompting. Recently, there has been a growing interest in transferring these reasoning abilities from LLMs to smaller models. However, achieving both the diversity and consistency…

Cited by 0SourcecodeScholar
2021

Directed Acyclic Graph Network for Conversational Emotion Recognition

ACL 2021long

The modeling of conversational context plays a vital role in emotion recognition from conversation (ERC). In this paper, we put forward a novel idea of encoding the utterances with a directed acyclic graph (DAG) to better model the intrinsic structure within a conversation, and design a directed acy…