← Search

S. Khatamifard

1 accepted papers

2024

LLM in a flash: Efficient Large Language Model Inference with Limited Memory

ACL 2024long

Large language models (LLMs) are central to modern natural language processing, delivering exceptional performance in various tasks. However, their substantial computational and memory requirements present challenges, especially for devices with limited DRAM capacity. This paper tackles the challeng…

Cited by 113SourcePDFScholar