ICML 2025poster0 citations

Layer-wise Quantization for Quantized Optimistic Dual Averaging

Anh Duc Nguyen, Ilia Markov, Zhengqing Wu, Ali Ramezani-Kebrya, Kimon Antonakopoulos, Dan Alistarh, Volkan Cevher

Abstract

Modern deep neural networks exhibit heterogeneity across numerous layers of various types such as residuals, multi-head attention, etc., due to varying structures (dimensions, activation functions, etc.), distinct representation characteristics, which impact predictions. We develop a general layer-wise quantization framework with tight variance and code-length bounds, adapting to the heterogeneities over the course of training. We then apply a new layer-wise quantization technique within distributed variational inequalities (VIs), proposing a novel Quantized Optimistic Dual Averaging (QODA) algorithm with adaptive learning rates, which achieves competitive convergence rates for monotone VIs. We empirically show that QODA achieves up to a $150$% speedup over the baselines in end-to-end training time for training Wasserstein GAN on $12+$ GPUs.

Adaptive CompressionLayer-wise CompressionOptimistic Dual AveragingDistributed Variational Inequality
BibTeX
@inproceedings{
nguyen2025layerwise,
title={Layer-wise Quantization for Quantized Optimistic Dual Averaging},
author={Anh Duc Nguyen and Ilia Markov and Zhengqing Wu and Ali Ramezani-Kebrya and Kimon Antonakopoulos and Dan Alistarh and Volkan Cevher},
booktitle={Forty-second International Conference on Machine Learning},
year={2025},
url={https://openreview.net/forum?id=J6LYjEOxbz}
}
Layer-wise Quantization for Quantized Optimistic Dual Averaging · ICML 2025