← Search

Ivan Chelombiev

2 accepted papers

2024

SparQ Attention: Bandwidth-Efficient LLM Inference

ICML 2024poster

The computational difficulties of large language model (LLM) inference remain a significant obstacle to their widespread deployment. The need for many applications to support long input sequences and process them in large batches typically causes token-generation to be bottlenecked by data transfer.…

2019

Adaptive Estimators Show Information Compression in Deep Neural Networks

ICLR 2019poster

To improve how neural networks function it is crucial to understand their learning process. The information bottleneck theory of deep learning proposes that neural networks achieve good generalization by compressing their representations to disregard information that is not relevant to the task. How…

Cited by 52SourcePDFScholar