2025
Achieving binary weight and activation for LLMs using Post-Training Quantization
ACL 2025finding
Quantizing large language models (LLMs) to 1-bit precision significantly reduces computational costs, but existing quantization techniques suffer from noticeable performance degradation when using weight and activation precisions below 4 bits (W4A4). In this paper, we propose a post-training quantiz…