2025
RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models
ICML 2025poster
Supervised fine-tuning is a standard method for adapting pre-trained large language models (LLMs) to downstream tasks. Quantization has been recently studied as a post-training technique for efficient LLM deployment. To obtain quantized fine-tuned LLMs, conventional pipelines would first fine-tune t…