← Search

Nilesh Malpeddi

1 accepted papers

2025

CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compression

EMNLP 2025

Large Language Models (LLMs) typically rely on a large number of parameters for token embedding, leading to substantial storage requirements and memory footprints. In particular, LLMs deployed on edge devices are memory-bound, and reducing the memory footprint by compressing the embedding layer not

Cited by 0SourcePDFScholar