← Search

Yadong Zhang

7 accepted papers

2026

SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs

CVPR 2026

Knowledge distillation (KD) is a standard route to compress Large Language Models (LLMs) into compact students, yet most pipelines uniformly apply token-wise loss regardless of teacher confidence. This indiscriminate supervision amplifies noisy, high-entropy signals and is especially harmful under l

Cited by 0SourcecodeScholar
2025

DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation Trainer

NeurIPS 2025poster

Recent advances in knowledge distillation have emphasized the importance of decoupling different knowledge components. While existing methods utilize momentum mechanisms to separate task-oriented and distillation gradients, they overlook the inherent conflict between target-class and non-target-clas…

Cited by 1SourcecodeScholar
2025

K-Level Reasoning: Establishing Higher Order Beliefs in Large Language Models for Strategic Reasoning

NAACL 2025long

Strategic reasoning is a complex yet essential capability for intelligent agents. It requires Large Language Model (LLM) agents to adapt their strategies dynamically in multi-agent environments. Unlike static reasoning tasks, success in these contexts depends on anticipating other agents’ beliefs an…

Cited by 0SourcePDFScholar
2024

Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models

NeurIPS 2024poster

Large language models (LLMs) have exhibited impressive performance in language comprehension and various reasoning tasks. However, their abilities in spatial reasoning, a crucial aspect of human cognition, remain relatively unexplored. Human possess a remarkable ability to create mental images of un…

2024

TOREE: Evaluating Topic Relevance of Student Essays for Chinese Primary and Middle School Education

ACL 2024findings

Topic relevance of an essay demands that the composition adheres to a clear theme and aligns well with the essay prompt requirements, a critical aspect of essay quality evaluation. However, existing research of Automatic Essay Scoring (AES) for Chinese essays has overlooked topic relevance and lacks…

Cited by 6SourcePDFScholar
2024

Towards Explainable Chinese Native Learner Essay Fluency Assessment: Dataset, Tasks, and Method

EMNLP 2024finding

Grammatical Error Correction (GEC) is a crucial technique in Automated Essay Scoring (AES) for evaluating the fluency of essays. However, in Chinese, existing GEC datasets often fail to consider the importance of specific grammatical error types within compositional scenarios, lack research on data…

2023

Connective Prediction for Implicit Discourse Relation Recognition via Knowledge Distillation

ACL 2023long

Implicit discourse relation recognition (IDRR) remains a challenging task in discourse analysis due to the absence of connectives. Most existing methods utilize one-hot labels as the sole optimization target, ignoring the internal association among connectives. Besides, these approaches spend lots o…

Cited by 13SourcePDFScholar