2026
Steering Pretrained Drafters During Speculative Decoding
AAAI 2026technical
Speculative decoding accelerates language model inference by separating generation into fast drafting and parallel verification. Its main limitation is drafter–verifier misalignment, which limits token acceptance and reduces overall effectiveness. While small drafting heads trained from scratch com