← Search

Zishun Yu

8 accepted papers

2026

Language Model Distillation: A Temporal Difference Imitation Learning Perspective

AAAI 2026technical

Large language models have led to significant progress across many NLP tasks, although their massive sizes often incur substantial computational costs. Distillation has become a common practice to compress these large and highly capable models into smaller, more efficient ones. Many existing languag

Cited by 0SourcePDFScholar
2025

Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization

ICML 2025poster

Solving mathematics problems has been an intriguing capability of large language models, and many efforts have been made to improve reasoning by extending reasoning length, such as through self-correction and extensive long chain-of-thoughts. While promising in problem-solving, advanced long reasoni…

Cited by 7SourcePDFScholar
2025

Towards Efficient Collaboration via Graph Modeling in Reinforcement Learning

AAAI 2025technical

In multi-agent reinforcement learning, a commonly considered paradigm is centralized training with decentralized execution. However, in this framework, decentralized execution restricts the development of coordinated policies due to the local observation limitation. In this paper, we consider the co…

Cited by 1SourcePDFScholar
2024

$\mathcal{B}$-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis

ICLR 2024spotlight

Program synthesis aims to create accurate, executable programs from problem specifications, specifically from natural language descriptions in our context. Recent studies have leveraged the power of reinforcement learning (RL) in conjunction with large language models (LLMs), significantly enhancin…

Cited by 2SourcePDFScholar
2022

Certifying Robust Graph Classification under Orthogonal Gromov-Wasserstein Threats

NeurIPS 2022accept

Graph classifiers are vulnerable to topological attacks. Although certificates of robustness have been recently developed, their threat model only counts local and global edge perturbations, which effectively ignores important graph structures such as isomorphism. To address this issue, we propose m…

Cited by 5SourcePDFScholar