2026
DyLLM: Efficient Diffusion LLM inference via saliency-based token selection and partial attention
ICML 2026poster
Masked Diffusion Language Models (MDLMs) enable parallel token decoding, providing a promising alternative to the sequential nature of autoregressive generation. However, their iterative denoising process remains computationally expensive because it repeatedly processes the entire sequence at every …