← Search

Sayeh Sharify

2 accepted papers

2025

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

ICML 2025spotlight

Post-training quantization (PTQ) of large language models (LLMs) holds the promise in reducing the prohibitive computational cost at inference time. Quantization of all weight, activation and key-value (KV) cache tensors to 4-bit without significantly degrading generalizability is challenging, due t…

2017

Bit-Pragmatic Deep Neural Network Computing

ICLR 2017workshop

We quantify a source of ineffectual computations when processing the multiplications of the convolutional layers in Deep Neural Networks (DNNs) and propose Pragrmatic (PRA), an architecture that exploits it improving performance and energy efficiency. The source of these ineffectual computations is…

Cited by 322SourceScholar