← Search

Sergei Krutikov

1 accepted papers

2026

LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding

ICML 2026poster

Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to propose candidate tokens that are then verified in parallel by the target model. The speedup is significantly determined by the acceptance rate, yet standard training minimizes …

Cited by 0SourceScholar