← Search

Asaf Aharoni

1 accepted papers

2025

Block Verification Accelerates Speculative Decoding

ICLR 2025poster

Speculative decoding is an effective method for lossless acceleration of large language models during inference. It uses a fast model to draft a block of tokens which are then verified in parallel by the target model, and provides a guarantee that the output is distributed identically to a sample f…

Cited by 3SourcePDFScholar