2026
Re-SpS: A Reinforcement Learning Approach to Speculative Sampling
AAAI 2026technical
Inference time latency has remained an open challenge for real world applications of large language models (LLMs). State-of-the-art (SOTA) speculative sampling (SpS) methods for LLMs, like EAGLE-3, use tree-based drafting to explore multiple candidate continuations in parallel. However, the hyperpar