2026
BPDQ: Bit-Plane Decomposition Quantization on a Variable Grid for Large Language Models
ICML 2026poster
Large language model (LLM) inference is often bounded by memory footprint and memory bandwidth in resource-constrained deployments, making quantization a fundamental technique for efficient serving. While post-training quantization (PTQ) maintains high fidelity at 4-bit, it deteriorates at 2–3 bits.…