FOCUS: DLLMs Know How to Tame Their Compute Bound
Kaihua Liang, Xin Tan, An Zhong, Hong Xu, Marco Canini
Abstract
Diffusion Large Language Models (**DLLMs**) offer a compelling alternative to Auto-Regressive models, but their deployment is constrained by high decoding cost. In this work, we identify a key inefficiency in DLLM decoding: while computation is parallelized over token blocks, only a small subset of tokens is decodable at each diffusion step, causing most compute to be wasted on non-decodable tokens. We further observe a strong correlation between attention-derived token importance and token-wise decoding probability. Based on this insight, we propose **FOCUS**—an inference system designed for DLLMs. By dynamically *focusing* computation on decodable tokens and evicting non-decodable ones on-the-fly, FOCUS increases the effective batch size, alleviating compute limitations and enabling scalable throughput. Empirical evaluations demonstrate that FOCUS achieves up to **3.52× throughput** improvement over the production-grade engine LMDeploy, while preserving or improving generation quality across multiple benchmarks.
BibTeX
@inproceedings{
liang2026focus,
title={{FOCUS}: {DLLM}s Know How to Tame Their Compute Bound},
author={Kaihua Liang and Xin Tan and An Zhong and Hong Xu and Marco Canini},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=40fUEdwvH3}
}