2024
QUIK: Towards End-to-end 4-Bit Inference on Generative Large Language Models
EMNLP 2024main
Large Language Models (LLMs) from the GPT family have become extremely popular, leading to a race towards reducing their inference costs to allow for efficient local computation. However, the vast majority of existing work focuses on weight-only quantization, which can reduce runtime costs in the me…