AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
The deployment of large language models (LLMs) is increasingly constrained by memory and latency bottlenecks, motivating the need for quantization techniques that flexibly balance accuracy and efficiency. Recent work has introduced multi-precision models, which enable inference at multiple precisio…