2026
ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
ICML 2026oral
Speculative Decodin promises to accelerate Large Language Model inference, yet its efficacy often degrades in production-grade scenarios. Existing evaluations typically overlook the compute-bound nature of high-concurrency regimes, where verification compute becomes the dominant bottleneck. Conseque…