← Search

Toby James Boyd

1 accepted papers

2024

Tandem Transformers for Inference Efficient LLMs

ICML 2024poster

The autoregressive nature of conventional large language models (LLMs) inherently limits inference speed, as tokens are generated sequentially. While speculative (Leviathan et al., 2023) and parallel (Stern et al., 2018) decoding techniques attempt to mitigate this, they face limitations: either rel…

Cited by 5SourcePDFScholar