← Search

Jeffrey T. H. Wong

2 accepted papers

2026

A3: an Analytical Low-Rank Approximation Framework for Attention

ICML 2026poster

Large language models have demonstrated remarkable performance; however, their massive parameter counts make deployment highly expensive. Low-rank approximation offers a promising compression solution, yet existing approaches have two main limitations: (1) They focus on minimizing the output error o…

Cited by 0SourceScholar
2025

QERA: an Analytical Framework for Quantization Error Reconstruction

ICLR 2025poster

The growing number of parameters and computational demands of large language models (LLMs) present significant challenges for their efficient deployment. Recently, there is an increasing interest in quantizing weights to extremely low precision while offsetting the resulting error with low-rank, hig…

Cited by 0SourcePDFScholar