FLEXI-LORA WITH INPUT-ADAPTIVE RANKS: EFFICIENT FINETUNING FOR SPEECH AND REASONING TASKS
Parameter-efficient fine-tuning methods like Low-Rank Adaptation (LoRA) have become essential for deploying large language models, yet their static parameter allocation remains suboptimal for inputs of varying complexity. We present Flexi-LoRA, a novel framework that dynamically adjusts LoRA ranks b…