← Search

Shaojie Zhuo

2 accepted papers

2025

OmniDraft: A cross-vocabulary, online adaptive drafter for on-device speculative decoding

NeurIPS 2025poster

Speculative decoding generally dictates having a small, efficient draft model that is either pretrained or distilled offline to a particular target model series, for instance, Llama or Qwen models. However, within online deployment settings, there are two major challenges: 1) usage of a target model…

Cited by 0SourceScholar
2024

Stepping Forward on the Last Mile

NeurIPS 2024poster

Continuously adapting pre-trained models to local data on resource constrained edge devices is the \emph{last mile} for model deployment. However, as models increase in size and depth, backpropagation requires a large amount of memory, which becomes prohibitive for edge devices. In addition, most ex…

Cited by 1SourcePDFScholar