2025
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
NeurIPS 2025poster
Quantization plays a crucial role in accelerating the inference of large-scale models, and rotational matrices have been shown to effectively improve quantization performance by smoothing outliers. However, end-to-end fine-tuning of rotational optimization algorithms incurs high computational costs…