← Search

Yineng Zhang

4 accepted papers

2026

Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving

ICML 2026poster

Prefill-Decode (PD) disaggregation has become the standard architecture for modern LLM inference engines, which alleviates the interference of two distinctive workloads. With the growing demand for multi-turn interactions in chatbots and agentic systems, we re-examined PD in this case and found two …

Cited by 0SourceScholar
2026

SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding

ICML 2026poster

Speculative decoding mitigates the memory-bound nature of LLM decoding by using a lightweight draft model to propose multiple tokens for parallel verification. However, its adoption has been limited by the lack of high-quality draft models and scalable training infrastructure. We introduce SpecForge…

Cited by 0SourceScholar
2026

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System

ICML 2026poster

Speculative decoding can significantly accelerate LLM serving, but its real-world benefits often erode due to training–serving mismatch and non-stationary traffic. Unlike previous systems that decouple speculator training from inference, we present a unified training–serving system, Aurora, that clo…

Cited by 0SourceScholar
2024

POSE-HMR: Heuristic Transformer with Postural Prior Constraints for 3D Human Mesh Reconstruction

ICASSP 2024accepted

This paper proposes an efficient and lightweight model called PoseHMR to address the interference of irrelevant image features and the issues of model inefficiency in 3D human body mesh reconstruction. PoseHMR uses a transformer-decoder architecture and obtains holistic and regional prior constraint…

Cited by 0SourceScholar