← Search

Ziyin Gu

4 accepted papers

2026

Group Causal Policy Optimization for Post-Training Large Language Models

AAAI 2026technical

Recent advances in large language models (LLMs) have broadened their applicability across diverse tasks, yet specialized domains still require targeted post-training. Among existing methods, Group Relative Policy Optimization (GRPO) stands out for its efficiency, leveraging groupwise relative reward

Cited by 0SourcePDFScholar
2026

On the Plasticity and Stability for Post-Training Large Language Models

ICML 2026poster

Training stability remains a critical bottleneck for Group Relative Policy Optimization (GRPO), often manifesting as a trade-off between reasoning plasticity and general capability retention. We identify a root cause as the geometric conflict between plasticity and stability gradients, which leads t…

Cited by 0SourceScholar
2025

Domain-Aware Knowledge Debiasing for Generalizable Video Understanding in CLIP

ICASSP 2025accepted

The pre-trained models contain multitudinous knowledge from huge amount of data. However, when applying these models to downstream tasks, they may mis-locate to wrong knowledge distribution due to a lack of domain or contextual knowledge. To address the distribution bias between the pre-trained mode…

Cited by 0SourceScholar
2025

Heterogeneous Packet Translation for Cross-Technology Communication

ICASSP 2025accepted

Recent advances in cross-technology communication (CTC) enable heterogeneous wireless devices (e.g., WiFi, Zig-Bee, and BLE) operating in the ISM band to communicate and understand each other. However, due to the limitation of standards and devices, existing CTC techniques need to design specific sc…

Cited by 0SourceScholar