← Search

Roberto Garcia

1 accepted papers

2025

Adaptive Rank Allocation: Speeding Up Modern Transformers with RaNA Adapters

ICLR 2025poster

Large Language Models (LLMs) are computationally intensive, particularly during inference. Neuron-adaptive techniques, which selectively activate neurons in Multi-Layer Perceptron (MLP) layers, offer some speedups but suffer from limitations in modern Transformers. These include reliance on sparse a…