2025
Surprising Effectiveness of pretraining Ternary Language Model at Scale
ICLR 2025spotlight
Rapid advancements in GPU computational power has outpaced memory capacity and bandwidth growth, creating bottlenecks in Large Language Model (LLM) inference. Post-training quantization is the leading method for addressing memory-related bottlenecks in LLM inference, but it suffers from significant…