← Search

Alexander Borzunov

5 accepted papers

2024

SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

ICLR 2024poster

Recent advances in large language model (LLM) pretraining have led to high-quality LLMs with impressive abilities. By compressing such LLMs via quantization to 3-4 bits per parameter, they can fit into memory-limited devices such as laptops and mobile phones, enabling personalized use. Quantizing mo…

2023

Distributed Inference and Fine-tuning of Large Language Models Over The Internet

NeurIPS 2023poster

Large language models (LLMs) are useful in many NLP tasks and become more capable with size, with the best open-source models having over 50 billion parameters. However, using these 50B+ models requires high-end hardware, making them inaccessible to most researchers. In this work, we investigate met…

Cited by 56SourcePDFScholar
2023

SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient

ICML 2023poster

Many deep learning applications benefit from using large models with billions of parameters. Training these models is notoriously expensive due to the need for specialized HPC clusters. In this work, we consider alternative setups for training large models: using cheap ``preemptible'' instances or p…

2022

Secure Distributed Training at Scale

ICML 2022spotlight

Many areas of deep learning benefit from using increasingly larger neural networks trained on public data, as is the case for pre-trained models for NLP and computer vision. Training such models requires a lot of computational resources (e.g., HPC clusters) that are not available to small research g…

2021

Distributed Deep Learning In Open Collaborations

NeurIPS 2021poster

Modern deep learning applications require increasingly more compute to train state-of-the-art models. To address this demand, large corporations and institutions use dedicated High-Performance Computing clusters, whose construction and maintenance are both environmentally costly and well beyond the…

Cited by 62SourcePDFScholar