← Search

Fangxin Liu

5 accepted papers

2026

SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization

AAAI 2026technical

The emergence of accurate open large language models (LLMs) has sparked a push for advanced quantization techniques to enable efficient deployment on end-user devices. In this paper, we revisit the challenge of extreme LLM compression---targeting ultra-low-bit quantization for both activations and w

Cited by 0SourcePDFScholar
2025

FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization

EMNLP 2025

The rapid advancement of large language models (LLMs) has exacerbated the memory bottleneck due to the widening gap between model parameter scaling and hardware capabilities. While post-training quantization techniques effectively reduce memory overhead, existing methods predominantly rely on static

Cited by 0SourcePDFScholar
2022

DynSNN: A Dynamic Approach to Reduce Redundancy in Spiking Neural Networks

ICASSP 2022accepted

Current Internet of Things (IoT) embedded applications use machine learning algorithms to process the collected data. However, the computational complexity and storage requirements of existing deep learning methods hinder the wide availability of embedded applications. Spiking Neural Networks (SNN)…

Cited by 0SourceScholar
2022

SpikeConverter: An Efficient Conversion Framework Zipping the Gap between Artificial Neural Networks and Spiking Neural Networks

AAAI 2022technical

Spiking Neural Networks (SNNs) have recently attracted enormous research interest since their event-driven and brain-inspired structure enables low-power computation. In image recognition tasks, the best results are achieved by SNN so far utilizing ANN-SNN conversion methods that replace activation…

Cited by 56SourcePDFScholar
2021

Improving Neural Network Efficiency via Post-Training Quantization With Adaptive Floating-Point

ICCV 2021poster

Model quantization has emerged as a mandatory technique for efficient inference with advanced Deep Neural Networks (DNN). It converts the model parameters in full precision (32-bit floating point) to the hardware friendly data representation with shorter bit-width, to not only reduce the model size…

Cited by 58PDFcodeScholar