← Search

Juncan Deng

5 accepted papers

2026

Rethinking Residual Errors in Compensation-based LLM Quantization

ICLR 2026poster

Methods based on weight compensation, which iteratively apply quantization and weight compensation to minimize the output error, have recently demonstrated remarkable success in quantizing Large Language Models (LLMs). The representative work, GPTQ, introduces several key techniques that make such…

Cited by 0SourcecodeScholar
2025

SSVQ: Unleashing the Potential of Vector Quantization with Sign-Splitting

ICCV 2025poster

Vector Quantization (VQ) has emerged as a prominent weight compression technique, showcasing substantially lower quantization errors than uniform quantization across diverse models, particularly in extreme compression scenarios. However, its efficacy during fine-tuning is limited by the constraint o…

2025

VQ4DiT: Efficient Post-Training Vector Quantization for Diffusion Transformers

AAAI 2025technical

The Diffusion Transformers Models (DiTs) have transitioned the network architecture from traditional UNets to transformers, demonstrating exceptional capabilities in image generation. Although DiTs have been widely applied to high-definition video generation tasks, their large parameter size hinders…

Cited by 8SourcePDFScholar
2025

ViM-VQ: Efficient Post-Training Vector Quantization for Visual Mamba

ICCV 2025poster

Visual Mamba networks (ViMs) extend the selective state space model (Mamba) to various vision tasks and demonstrate significant potential. As a promising compression technique, vector quantization (VQ) decomposes network weights into codebooks and assignments, significantly reducing memory usage and…

Cited by 0SourcePDFScholar