2025
zFLoRA: Zero-Latency Fused Low-Rank Adapters
EMNLP 2025
Large language models (LLMs) are increasingly deployed with task-specific adapters catering to multiple downstream applications. In such a scenario, the additional compute associated with these apparently insignificant number of adapter parameters (typically less than 1% of the base model) turns out