← Search

Zhangming Li

2 accepted papers

2026

MemeBQ:Memory Efficient Binary Quantization of LLMs

AAAI 2026technical

Recent years have witnessed growing scholarly interest in binary post-training quantization (PTQ) techniques for large language models (LLMs). While state-of-the-art (SOTA) binary quantization methods significantly reduce memory footprint and computational demands, they introduce additional memory o

Cited by 0SourcePDFScholar
2025

LoRaDA: Low-Rank Direct Attention Adaptation for Efficient LLM Fine-tuning

EMNLP 2025

As the parameter size of language models becomes extremely large, fine-tuning them with limited resources has become a challenging task. Latest advancements in parameter-efficient fine-tuning (PEFT) techniques allow for adjustments to only a minor fraction of the parameters of these LLMs. Yet, most

Cited by 0SourcePDFScholar