2026
CoCoQuant: Breaking the Bandwidth Wall via Co-Optimized Communication and Computation Quantization
ICML 2026poster
The rapid scaling of large language models (LLMs) has made distributed inference indispensable, yet end-to-end latency is increasingly dominated by communication, forming a critical bandwidth wall that fundamentally limits the practical gains of existing quantization techniques. Existing approaches …