← Search

Wenya Yu

1 accepted papers

2026

ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression

ICML 2026poster

Post-Training Quantization (PTQ) and Low-Rank Adaptation (LoRA) constitute the standard pipeline for efficient Large Language Model (LLM) deployment. However, applying them sequentially poses a problem: PTQ often leaves behind random noise that is spread out (across the model's weights) in a way LoR…

Cited by 0SourcecodeScholar