← Search

Yize Wu

3 accepted papers

2025

EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU Utilization

NeurIPS 2025poster

Speculative decoding is an effective and lossless method for Large Language Model (LLM) inference acceleration. It employs a smaller model to generate a draft token sequence, which is then verified by the original base model. In multi-GPU systems, inference latency can be further reduced through ten…

Cited by 0SourcecodeScholar