2026
MemeBQ:Memory Efficient Binary Quantization of LLMs
AAAI 2026technical
Recent years have witnessed growing scholarly interest in binary post-training quantization (PTQ) techniques for large language models (LLMs). While state-of-the-art (SOTA) binary quantization methods significantly reduce memory footprint and computational demands, they introduce additional memory o