← Search

Luohe Shi

1 accepted papers

2026

Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios

AAAI 2026technical

Speculative decoding accelerates LLM inference by utilizing otherwise idle computational resources during memory-to-chip data transfer. Current speculative decoding methods typically assume a considerable amount of available computing power, then generate a complex and massive draft tree using a sma

Cited by 0SourcePDFScholar