← Search

Justus Will

2 accepted papers

2026

Parallel Token Generation for Language Models

ICLR 2026poster

Autoregressive transformers are the backbone of modern large language models. Despite their success, inference remains slow due to strictly sequential prediction. Prior attempts to predict multiple tokens per step typically impose independence assumptions across tokens, which limits their ability to…

Cited by 0SourceScholar