← Search

Lijuan Hu

2 accepted papers

2026

FLRQ: Faster LLM Quantization with Flexible Low-Rank Matrix Sketching

AAAI 2026technical

Traditional post-training quantization (PTQ) is considered an effective approach to reduce model size and accelerate inference of large-scale language models (LLMs). However, existing low-rank PTQ methods require costly fine-tuning to determine a compromise rank for diverse data and layers in large

Cited by 0SourcePDFScholar
2026

TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling

ICML 2026poster

Mixture-of-Experts (MoE) models achieve remarkable performance by sparsely activating specialized experts, yet their massive parameters in experts pose significant challenges for deployment. While low-rank quantization offers a promising route to compress MoE models, existing methods still incur non…

Cited by 0SourceScholar