CAR-LoRA: Training Compression-Aware and Robust LoRA Adapters for Evolving LLMs
The deployment of large language models (LLMs) for specialized tasks on resource-constrained edge devices like smartphones and sensors presents a significant scalability problem. To run on such hardware, these massive models must be compressed using techniques like \emph{quantization or pruning} to…