← Search

Lancheng Zou

3 accepted papers

2026

S-Quant: Rethinking Weight Quantization with Seed-Based Generation

ICML 2026poster

The progressive scaling of large language models (LLMs) has consistently enhanced multimodal understanding and advanced reasoning capabilities, but has substantially increased computational and hardware execution overhead. In this paper, we present S-Quant, a novel post-method that compresses only m…

Cited by 0SourceScholar
2025

PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models

NeurIPS 2025poster

Channel permutation is a powerful technique for enhancing the accuracy of N:M sparse models by reordering the channels of weight matrices to prioritize the retention of important weights. However, traditional channel permutation methods rely on handcrafted quality metrics, which often fail to accur…

Cited by 0SourceScholar
2024

BiE: Bi-Exponent Block Floating-Point for Large Language Models Quantization

ICML 2024poster

Nowadays, Large Language Models (LLMs) mostly possess billions of parameters, bringing significant challenges to hardware platforms. Although quantization is an efficient approach to reduce computation and memory overhead for inference optimization, we stress the challenge that mainstream low-bit qu…

Cited by 5SourcePDFScholar