2026
ATTS: Asynchronous Test-Time Scaling via Conformal Prediction
ICLR 2026poster
Large language models (LLMs) benefit from test-time scaling but are often hampered by high inference latency. Speculative decoding is a natural way to accelerate the scaling process; however, scaling along both the parallel and sequential dimensions poses significant challenges, including substantia…