← Search

Nguyen-Khang Le

2 accepted papers

2026

AdaSpec: Adaptive Multilingual Speculative Decoding with Self-Synthesized Language-Aware Training and Vocabulary Simplification

AAAI 2026technical

Speculative decoding accelerates large language model (LLM) inference by using a lightweight drafter to propose multiple tokens, which are then verified in parallel by the base model. While effective in English, existing methods often struggle in multilingual scenarios due to static vocabularies and

Cited by 0SourcePDFScholar
2025

SPECTRA: Faster Large Language Model Inference with Optimized Internal and External Speculation

ACL 2025long

Inference with modern Large Language Models (LLMs) is both computationally expensive and time-consuming. Speculative decoding has emerged as a promising solution, but existing approaches face key limitations: training-based methods require a draft model that is challenging to obtain and lacks genera…

Cited by 0SourcePDFScholar