2024
SparQ Attention: Bandwidth-Efficient LLM Inference
ICML 2024poster
The computational difficulties of large language model (LLM) inference remain a significant obstacle to their widespread deployment. The need for many applications to support long input sequences and process them in large batches typically causes token-generation to be bottlenecked by data transfer.…