← Search

Shiqiang Nie

1 accepted papers

2026

KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache

AAAI 2026technical

The high memory demands of the Key-Value (KV) Cache during the inference of Large Language Models (LLMs) severely restrict their deployment in resource-constrained platforms. Quantization can effectively alleviate the memory pressure caused by KV Cache. However, existing methods either rely on stati

Cited by 0SourcePDFScholar