2025
CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compression
EMNLP 2025
Large Language Models (LLMs) typically rely on a large number of parameters for token embedding, leading to substantial storage requirements and memory footprints. In particular, LLMs deployed on edge devices are memory-bound, and reducing the memory footprint by compressing the embedding layer not