2025
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
EMNLP 2025
The rapid advancement of large language models (LLMs) has exacerbated the memory bottleneck due to the widening gap between model parameter scaling and hardware capabilities. While post-training quantization techniques effectively reduce memory overhead, existing methods predominantly rely on static