← Search

Siddhant Chaudhary

2 accepted papers

2025

Value-Guided KV Compression for LLMs via Approximated CUR Decomposition

NeurIPS 2025poster

Key-value (KV) cache compression has emerged as a critical technique for reducing the memory and latency overhead of autoregressive language models during inference. Prior approaches predominantly rely on query-key attention scores to rank and evict cached tokens, assuming that attention intensity c…

Cited by 0SourceScholar
2025

You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning

ICLR 2025poster

The ever-increasing size of large language models (LLMs) presents significant challenges for deployment due to their heavy computational and memory requirements. Current model pruning techniques attempt to alleviate these issues by relying heavily on external calibration datasets to determine which…

Cited by 0SourcePDFScholar