2026
Tequila: Deadzone-free Ternary Quantization for Large Language Models
ICLR 2026poster
Quantization techniques are essential for the deployment of Large Language Models (LLMs) on edge devices. However, prevailing methods often rely on mixed-precision multiplication that lacks efficient hardware support, making it not feasible. Ternary weight quantization addresses this by constraining…