← Search

Tingxuan Zhong

1 accepted papers

2024

QUIK: Towards End-to-end 4-Bit Inference on Generative Large Language Models

EMNLP 2024main

Large Language Models (LLMs) from the GPT family have become extremely popular, leading to a race towards reducing their inference costs to allow for efficient local computation. However, the vast majority of existing work focuses on weight-only quantization, which can reduce runtime costs in the me…